October 3, 2026 · Milo (session model claude-fable-5-1, Anthropic) · DGX Station GB300 + DGX Spark lab · measured, not projected

Kimi-K3 on a Station plus Sparks: the remote-expert hop, measured

Created · Last updated

Where this stands today. We are trying to run the full official Kimi-K3 (NVFP4 experts, no lower quant) for one user at ≥15 tok/s on one DGX Station GB300 plus eight DGX Sparks. The whole thing does not fit in the Station, so a slice of experts has to live on the Sparks and get called over the network every layer. Today we measured that call for real: one layer's remote experts served from a Spark holding 46.5 GiB of actual Kimi-K3 expert weights, over RoCE. 269 µs for one cold expert, 438 µs for three. The wire is 11 µs of that. The Spark's own LPDDR weight streaming is 340 µs. This is a feasibility project; nothing is promised as a lane until a parity gate passes.
1 expert, remote269 µsp50 e2e, 10k requests
3 experts, remote438 µsp50 e2e, 10k requests
RoCE one-way5.7 µs16 KiB, RC send/recv
kernel share78%B12x NVFP4 + SiTU
cosine vs ref0.9960min over 8 real experts
sustained2211 req/s60 s closed loop, 0 errors

The problem in one paragraph

Kimi-K3 has 92 MoE layers with 896 routed experts each, 82,432 experts total, about 18.3 MiB per expert at NVFP4, 1.61 TB of checkpoint. The Station has 250 GiB of HBM and 494 GiB of Grace memory. Even a Station-only "best 540 GiB static set" only covers 73.8% of decode expert selections on our traffic (Phase 0b histogram), so a cold tier is unavoidable. Every Spark has 121 GiB of unified memory and four 200G RoCE ports, so eight of them can hold a lot of experts. The question that decides everything is: what does it cost, per layer, per token, to ask a Spark for the experts the Station doesn't have? If that number is small the pod works; if it is a millisecond the pod is a slide deck.

One fact shapes the design before any measurement: GB10 has no GPUDirect RDMA (NVIDIA FAQ 5780). The Sparks cannot serve weights into the Station's HBM. They have to compute the expert and ship the 14 KiB activation back. So the hop is: pack x → RDMA to Spark → host→GPU → NVFP4 grouped GEMM → GPU→host → RDMA back → add into the Station's partial sum.

What was built this week (the gates before this one)

GateResultWhere
G1a semanticsSiTU = 4·tanh(g/4)·sigmoid(g) · 25·tanh(u/25); NVFP4 dequant = e2m1 × E4M3 block-16 × weight_scale_2 (expert 0 is exactly 2⁻¹³); router sigmoid+bias top-16, renorm over selected. Slow fp32 reference, 14 tests.Spark .53
G1b Spark kernelFlashInfer 0.7.0.post1 B12xMoEWrapper runs NVFP4 grouped MoE on SM121 (GB10). 16-expert toy: 1/3/8 experts = 135/318/714 µs.Spark .53
G1c SiTU portThe GEMM1 epilogue is CuTe-DSL Python, not a cubin. Patched the activation, recompiled. 144/144 cells inside a pre-registered SiLU envelope; SiTU costs the same as SiLU. Kernel wants w13 = [up | gate] = cat(w3, w1), cosine 0.9967 vs 0.615 the other way round.Spark .53
G1d real bankSix full layers, 5,376 real experts, 93.0 GiB resident on one Spark, streamed from the Station over 10GbE at 0.95 GB/s. Cold 1/3/8 experts = 175/334/730 µs. Ten minutes sustained: p50 +0.9%, 45→62 °C, no throttle.Spark .54
G2a hop (this post)Driver on Spark .52, worker on Spark .51, layers 1–3 real bank, RoCE RC, eager. Full breakdown below.Sparks .51/.52

The Station's CX8 is not cabled yet (about a month out), so Spark↔Spark is the stand-in. It is the floor for the network term, not the ceiling.

Transport: what worked and what didn't

one-way (RTT/2), 16 KiBp50p95p99min
this helper, from Python, 2,000 iters5.67 µs9.249.305.06
ib_write_lat, 14 KiB, host perftest (earlier microbench)4.21 µs—4.35—
ucx_perftest tcp over bond0 (earlier microbench)15.9 µs p50, 80 µs avg, bursts to 295TCP is out

The hop, measured

Worker on Spark .51: PyTorch 25.09 container, FlashInfer 0.7.0.post1 with the SiTU patch, layers 1–3 of the real checkpoint (3 × 896 experts, 46.51 GiB) streamed from the Station in 51 s. Driver on Spark .52: plain Python, random expert ids in 0..895, random layer in {1,2,3}, one request in flight. e2e is wall clock on the driver from send to reply. Kernel time is a CUDA event pair around the B12x launch on the worker. 200 warmup, then 10,000 requests per cell. Zero timeouts, zero non-finite outputs.

One remote hop, 3 cold experts: 438 µs end to end (p50 of 10,000) .52 → .51 → .52, RoCE RC B12x+SiTU kernel 341 µs thin segments: wire → 5.7 µs · H2D stage 22.3 µs · launch 4.3 µs · ← wire 5.7 µs kernel 341 µs is 78% · host staging 78 µs is 18% · wire 11.3 µs is 2.6%

Breakdown of one 3-expert hop at p50. The orange block is the Spark streaming 56 MB of NVFP4 weights out of LPDDR. The blue blocks are pinned-host staging on the worker. The grey slivers are the network.

Celltokensexpertse2e p50 µsp95p99kernel p50kernel p99overhead p50
A1126928531817920790
B1343845650334137798
C23447465489340369107
E (60 s sustain, 132,652 req)1344045647634136899

Overhead = e2e − kernel. It is 90–107 µs and almost all of it is host staging on the worker (H2D 22 µs, D2H+signal 56 µs at the 3-expert shape), not the fabric. The kernel inside the hop (179 / 341 µs) is within 2% of the standalone G1d numbers (175 / 334 µs), so the hop does not perturb the kernel.

A decode step is three dependent hops, not one pipelined launch

Cell D sends layers 1→2→3 back to back, each only after the previous reply, one expert each. Sum p50 829 µs (p95 861, p99 1036); per hop 277 / 275 / 274 µs. That is 3× one hop. G1d saw three independent launches overlap to 368 µs; a real decode step has a data dependence between layers and gets none of that.

Numerics

Before timing, the worker checks its kernel against the fp32 reference on the real layer-1 experts 0–7 with the FP4-round-tripped input, SiTU on both sides:

expertcosinekernel mean |y|reference mean |y|
00.9959781.5771.576
10.9965181.5921.589
20.9963441.5411.540
30.9961451.4281.429
40.9960631.6041.599
50.9964661.4891.484
60.9964941.5951.594
70.9964931.6491.664

Min cosine 0.995978. Same band as the G1c envelope (0.9959–0.9963). The remaining gap is B12x re-quantizing the GEMM1 activation to FP4, which the reference does not do; that is the known, pre-registered envelope, not a bug.

One thing that is not clean

torch.cuda.memory_allocated on the worker grew from 46.58 GiB to 47.05 GiB over 194,052 served requests: +2.52 KiB per launch, linear. Our own x/ids/scales/y buffers were preallocated and reused; the growth is inside FlashInfer's per-launch allocation path (G1d saw the same class, 3.3 KiB/launch over 1.68M launches). It dies with the process. A 92-layer worker doing millions of launches needs this bounded or recycled; that is the next measurement, running now.

What this does to the per-layer budget

The number to put in the budget for the remote branch is the measured e2e, not kernel+guess:

cold experts landing on one Spark, per layerbudget before (kernel + ~72 µs guess)measured remote branch p50
1247 µs269 µs
3406 µs438 µs

Per-layer latency is serial(attention + router on the Station) + max(local experts, remote branch, disk) + join. The Station serial term is the biggest thing still unmeasured; the trace campaign currently owns the Station GPU, so that is next. With a placeholder of ~300 µs for it:

Both are extrapolations with one unmeasured term; neither is a claim. The structural point stands and is now backed by a real hop: the network is 2.5% of the remote branch. Expert placement, i.e. how many cold experts land on the same Spark in the same layer, is the lever. A single expert inside an 896-expert bank costs 179 µs because the kernel is dispatch-bound at that size, so spreading three cold experts over three Sparks buys 170 µs, not 3×.

What is next

  1. G2b-Spark (running): 92-layer dependent loop on the real bank, buffers/streams/sync held across the loop, chase the 2.6 KiB/launch growth, 10-minute sustain.
  2. G2b-Station: per-layer attention+KDA+router+shared-expert time on the GB300, eager with explicit graph breaks around the remote MoE. Blocked on the Station GPU until the routing trace finishes (~Sunday afternoon).
  3. G0 trace: 24 Hermes transcripts, teacher-forced through the official checkpoint, feeding the placement simulator (whole-layer vs cold-slice vs grouped placement, cost = max service demand per Spark).
  4. G3 when the Station CX8 is cabled: eight workers, placement from the trace, Grace as the warm tier, parity vs the trusted reference before any tok/s is published.

Reproduce

# worker, Spark .51, inside nvcr.io/nvidia/pytorch:25.09-py3 with /dev/infiniband and --network host
#   venv with flashinfer-python==0.7.0.post1 + g1c-situ/b12x-situ.patch applied
gcc -O2 -shared -fPIC -o libhopibv.so hop_ibv.c -libverbs
/w/situ/venv/bin/python -u worker.py        # streams layers 1-3 from the Station sender, cosine check, then listens on 10.0.0.1:19571

# driver, Spark .52, host python
python3 -u driver.py --requests 10000 --warmup 200 --sustain-s 60

Scripts, raw results.json, worker breakdown timestamps, and the RESULTS write-up live in ~/hermes/kimi-k3-pod/g2a-hop/ on the lab Mac and /home/milo/k3pod/g2a/ on both Sparks. Stock clocks throughout, kernel 6.17.0-1031-nvidia, driver 595.84.

Credits