DGX SPARK
Connecting Manage access

FOUR-SPARK BLACKWELL INFERENCE

Your private frontier
model cluster.

One model-aware endpoint backed by four DGX Sparks and a switched 200 GbE ConnectX-7 fabric. Live telemetry is public; inference requires an API key.

NOW SERVING

Discovering model…

Waiting for the inference backend.

https://dgx.cautela.io/v1

LIVE TELEMETRY

Performance at a glance

Not sampled yet

Generationtokens / second
Prompt processingtokens / second
Active requests— queued
KV cacheallocated

ACCELERATORS

Cluster load

4 × GB10

Waiting for GPU telemetry…

LIFETIME

Serving counters

Completed requests
Generated tokens
Prompt tokens
Prefix cache hit rate
Average first token
Preemptions

GATEWAY

Request pressure

Since restart
Active requests
Streaming / buffered
Oldest active
Admission queue
Overload responses
Client disconnects
Maximum header wait
Last 256 ≥ 120 s

INTRA-DGX FABRIC

Switch traffic

Connecting
Bidirectional traffic 200 Gb/s each direction
Spark 01 · Head →→ Spark 02 · Worker
0% — utilized
Spark 01 · Head ←← Spark 02 · Worker
0% — utilized
Head link
Worker link
Spark 03 link
Spark 04 link
FEC corrected
FEC uncorrected
RX errors
TX drops

GET CONNECTED

Use the cluster from your tools

Ask the operator for a personal dgx_… key. Keep it in an environment variable and never commit it.

01

Launch Claude Code natively

Claude Code talks directly to the server's Anthropic-compatible Messages endpoint. No LiteLLM bridge is involved.

Loading current model…

Requires the claude CLI. The model aliases ensure every Claude Code agent role uses the active DGX model.