Kimi-K3 on a Station plus Sparks: the remote-expert hop, measured
The problem in one paragraph
Kimi-K3 has 92 MoE layers with 896 routed experts each, 82,432 experts total, about 18.3 MiB per expert at NVFP4, 1.61 TB of checkpoint. The Station has 250 GiB of HBM and 494 GiB of Grace memory. Even a Station-only "best 540 GiB static set" only covers 73.8% of decode expert selections on our traffic (Phase 0b histogram), so a cold tier is unavoidable. Every Spark has 121 GiB of unified memory and four 200G RoCE ports, so eight of them can hold a lot of experts. The question that decides everything is: what does it cost, per layer, per token, to ask a Spark for the experts the Station doesn't have? If that number is small the pod works; if it is a millisecond the pod is a slide deck.
One fact shapes the design before any measurement: GB10 has no GPUDirect RDMA (NVIDIA FAQ 5780). The Sparks cannot serve weights into the Station's HBM. They have to compute the expert and ship the 14 KiB activation back. So the hop is: pack x → RDMA to Spark → host→GPU → NVFP4 grouped GEMM → GPU→host → RDMA back → add into the Station's partial sum.
What was built this week (the gates before this one)
| Gate | Result | Where |
|---|---|---|
| G1a semantics | SiTU = 4·tanh(g/4)·sigmoid(g) · 25·tanh(u/25); NVFP4 dequant = e2m1 × E4M3 block-16 × weight_scale_2 (expert 0 is exactly 2⁻¹³); router sigmoid+bias top-16, renorm over selected. Slow fp32 reference, 14 tests. | Spark .53 |
| G1b Spark kernel | FlashInfer 0.7.0.post1 B12xMoEWrapper runs NVFP4 grouped MoE on SM121 (GB10). 16-expert toy: 1/3/8 experts = 135/318/714 µs. | Spark .53 |
| G1c SiTU port | The GEMM1 epilogue is CuTe-DSL Python, not a cubin. Patched the activation, recompiled. 144/144 cells inside a pre-registered SiLU envelope; SiTU costs the same as SiLU. Kernel wants w13 = [up | gate] = cat(w3, w1), cosine 0.9967 vs 0.615 the other way round. | Spark .53 |
| G1d real bank | Six full layers, 5,376 real experts, 93.0 GiB resident on one Spark, streamed from the Station over 10GbE at 0.95 GB/s. Cold 1/3/8 experts = 175/334/730 µs. Ten minutes sustained: p50 +0.9%, 45→62 °C, no throttle. | Spark .54 |
| G2a hop (this post) | Driver on Spark .52, worker on Spark .51, layers 1–3 real bank, RoCE RC, eager. Full breakdown below. | Sparks .51/.52 |
The Station's CX8 is not cabled yet (about a month out), so Spark↔Spark is the stand-in. It is the floor for the network term, not the ceiling.
Transport: what worked and what didn't
- ucxx-cu13 0.51.1 installs on aarch64, but its bundled UCX 1.20.1 ships no
rc/InfiniBand plugin.UCX_TLS=rcraisesUCXXNoDeviceError; the only transports it offers are tcp, cma, cuda, shm. Rejected. (Host UCX 1.16 has rc; the wheel does not.) - pyverbs opens the HCA fine (port ACTIVE, MTU 4096). We did not write the ring in it.
- Chosen: ~400 lines of libibverbs C wrapped with ctypes. RC queue pair on
rocep1s0f0GID index 3 (RoCEv2 IPv4), QP handshake over TCP, then pure SEND/RECV on eight registered 16 KiB slots. Slot 0 is send-only so a send and a receive are never armed on the same buffer; every send waits for its completion before the slot is reused; request ids are echoed and checked; polls time out instead of hanging. No NCCL.
| one-way (RTT/2), 16 KiB | p50 | p95 | p99 | min |
|---|---|---|---|---|
| this helper, from Python, 2,000 iters | 5.67 µs | 9.24 | 9.30 | 5.06 |
ib_write_lat, 14 KiB, host perftest (earlier microbench) | 4.21 µs | — | 4.35 | — |
ucx_perftest tcp over bond0 (earlier microbench) | 15.9 µs p50, 80 µs avg, bursts to 295 | TCP is out | ||
The hop, measured
Worker on Spark .51: PyTorch 25.09 container, FlashInfer 0.7.0.post1 with the SiTU patch, layers 1–3 of the real checkpoint (3 × 896 experts, 46.51 GiB) streamed from the Station in 51 s. Driver on Spark .52: plain Python, random expert ids in 0..895, random layer in {1,2,3}, one request in flight. e2e is wall clock on the driver from send to reply. Kernel time is a CUDA event pair around the B12x launch on the worker. 200 warmup, then 10,000 requests per cell. Zero timeouts, zero non-finite outputs.
Breakdown of one 3-expert hop at p50. The orange block is the Spark streaming 56 MB of NVFP4 weights out of LPDDR. The blue blocks are pinned-host staging on the worker. The grey slivers are the network.
| Cell | tokens | experts | e2e p50 µs | p95 | p99 | kernel p50 | kernel p99 | overhead p50 |
|---|---|---|---|---|---|---|---|---|
| A | 1 | 1 | 269 | 285 | 318 | 179 | 207 | 90 |
| B | 1 | 3 | 438 | 456 | 503 | 341 | 377 | 98 |
| C | 2 | 3 | 447 | 465 | 489 | 340 | 369 | 107 |
| E (60 s sustain, 132,652 req) | 1 | 3 | 440 | 456 | 476 | 341 | 368 | 99 |
Overhead = e2e − kernel. It is 90–107 µs and almost all of it is host staging on the worker (H2D 22 µs, D2H+signal 56 µs at the 3-expert shape), not the fabric. The kernel inside the hop (179 / 341 µs) is within 2% of the standalone G1d numbers (175 / 334 µs), so the hop does not perturb the kernel.
A decode step is three dependent hops, not one pipelined launch
Cell D sends layers 1→2→3 back to back, each only after the previous reply, one expert each. Sum p50 829 µs (p95 861, p99 1036); per hop 277 / 275 / 274 µs. That is 3× one hop. G1d saw three independent launches overlap to 368 µs; a real decode step has a data dependence between layers and gets none of that.
Numerics
Before timing, the worker checks its kernel against the fp32 reference on the real layer-1 experts 0–7 with the FP4-round-tripped input, SiTU on both sides:
| expert | cosine | kernel mean |y| | reference mean |y| |
|---|---|---|---|
| 0 | 0.995978 | 1.577 | 1.576 |
| 1 | 0.996518 | 1.592 | 1.589 |
| 2 | 0.996344 | 1.541 | 1.540 |
| 3 | 0.996145 | 1.428 | 1.429 |
| 4 | 0.996063 | 1.604 | 1.599 |
| 5 | 0.996466 | 1.489 | 1.484 |
| 6 | 0.996494 | 1.595 | 1.594 |
| 7 | 0.996493 | 1.649 | 1.664 |
Min cosine 0.995978. Same band as the G1c envelope (0.9959–0.9963). The remaining gap is B12x re-quantizing the GEMM1 activation to FP4, which the reference does not do; that is the known, pre-registered envelope, not a bug.
One thing that is not clean
torch.cuda.memory_allocated on the worker grew from 46.58 GiB to 47.05 GiB over 194,052 served requests: +2.52 KiB per launch, linear. Our own x/ids/scales/y buffers were preallocated and reused; the growth is inside FlashInfer's per-launch allocation path (G1d saw the same class, 3.3 KiB/launch over 1.68M launches). It dies with the process. A 92-layer worker doing millions of launches needs this bounded or recycled; that is the next measurement, running now.
What this does to the per-layer budget
The number to put in the budget for the remote branch is the measured e2e, not kernel+guess:
| cold experts landing on one Spark, per layer | budget before (kernel + ~72 µs guess) | measured remote branch p50 |
|---|---|---|
| 1 | 247 µs | 269 µs |
| 3 | 406 µs | 438 µs |
Per-layer latency is serial(attention + router on the Station) + max(local experts, remote branch, disk) + join. The Station serial term is the biggest thing still unmeasured; the trace campaign currently owns the Station GPU, so that is next. With a placeholder of ~300 µs for it:
- 3 cold experts on one Spark: ~300 + 438 ≈ 740 µs/layer × 92 ≈ 68 ms → ~14.7 tok/s.
- 1 cold expert per Spark, 3 Sparks in parallel: ~300 + 269 ≈ 570 µs/layer × 92 ≈ 52 ms → ~19 tok/s.
Both are extrapolations with one unmeasured term; neither is a claim. The structural point stands and is now backed by a real hop: the network is 2.5% of the remote branch. Expert placement, i.e. how many cold experts land on the same Spark in the same layer, is the lever. A single expert inside an 896-expert bank costs 179 µs because the kernel is dispatch-bound at that size, so spreading three cold experts over three Sparks buys 170 µs, not 3×.
What is next
- G2b-Spark (running): 92-layer dependent loop on the real bank, buffers/streams/sync held across the loop, chase the 2.6 KiB/launch growth, 10-minute sustain.
- G2b-Station: per-layer attention+KDA+router+shared-expert time on the GB300, eager with explicit graph breaks around the remote MoE. Blocked on the Station GPU until the routing trace finishes (~Sunday afternoon).
- G0 trace: 24 Hermes transcripts, teacher-forced through the official checkpoint, feeding the placement simulator (whole-layer vs cold-slice vs grouped placement, cost = max service demand per Spark).
- G3 when the Station CX8 is cabled: eight workers, placement from the trace, Grace as the warm tier, parity vs the trusted reference before any tok/s is published.
Reproduce
# worker, Spark .51, inside nvcr.io/nvidia/pytorch:25.09-py3 with /dev/infiniband and --network host
# venv with flashinfer-python==0.7.0.post1 + g1c-situ/b12x-situ.patch applied
gcc -O2 -shared -fPIC -o libhopibv.so hop_ibv.c -libverbs
/w/situ/venv/bin/python -u worker.py # streams layers 1-3 from the Station sender, cosine check, then listens on 10.0.0.1:19571
# driver, Spark .52, host python
python3 -u driver.py --requests 10000 --warmup 200 --sustain-s 60
Scripts, raw results.json, worker breakdown timestamps, and the RESULTS write-up live in ~/hermes/kimi-k3-pod/g2a-hop/ on the lab Mac and /home/milo/k3pod/g2a/ on both Sparks. Stock clocks throughout, kernel 6.17.0-1031-nvidia, driver 595.84.
Credits
- Moonshot AI — Kimi-K3; NVIDIA — the NVFP4 checkpoint, GB10/GB300, FlashInfer's B12x SM12x kernel path.
- FlashInfer — the CuTe-DSL micro kernel whose epilogue was patchable in Python.
- The Astra review that reordered the gates so this hop got measured before anything was integrated.