Switching 8 DGX Sparks and a GB300 Station: CRS804 ×2 vs one 32×400G switch

Created Last updated
Decision note · James Meadlock & Milo (James's AI agent) · written on claude-fable-5-1 (Anthropic), default reasoning · prompted by a vendor reply from NADDOD
Where this stands: the fabric is eight DGX Sparks at 200 GbE each plus one GB300 Station with two 400 GbE ports. One MikroTik CRS804 has exactly eight 200G lanes and nothing left for the Station. NADDOD's answer was "buy our N9200-32DC" and "two CRS804s add unavoidable inter-switch latency." Both statements are true. The latency is about 1 µs per extra hop, which is under 0.5% of a decode step. The real reasons to prefer one switch are trunk bandwidth, zero spare ports, and the fact that the Station wants two uplinks. The real reasons to hesitate are $10,459, a late-October batch, six datacenter fans, and 200–240 VAC input only. No order has changed yet.
~1 µs
added one-way latency per extra switch hop estimate
<0.5%
of a 30–50 ms decode step, worst-case TP=8 estimate
2:1
trunk oversubscription, Option A
$10,459
N9200-32DC list, 1-month lead
200–240 V
N9200 AC input; no 120 V

Contents

  1. The port problem
  2. The three options, drawn
  3. Quantifying "unavoidable inter-switch latency"
  4. The actual Option A problem: the trunk
  5. Price, power, noise, gotchas
  6. What the vendor got right and wrong
  7. What happens next

The port problem

Each DGX Spark exposes one ConnectX-7 port at 200 GbE. The GB300 Station has a ConnectX-8 with two QSFP112 cages at 400 GbE, running in Ethernet mode. The CRS804-4DDQ-hRM has four QSFP56-DD cages at 400G; each one breaks out to 2×200G with a Q2Q56 splitter DAC. Four cages × 2 = eight 200G lanes. Eight Sparks. Full. The Station has nowhere to plug in.

Until recently that was fine. The Sparks talked to each other and the Station was a separate island on 10 GbE. The plan changed once it became clear that a single model can usefully span HBM on the Station, Grace memory on the Station, and the Spark cluster together — a Kimi K3 NVFP4 style placement — and at that point the Station needs to be on the same fabric at a rate that is not 10 GbE.

The three options, drawn

Three fabric options for eight Sparks and one Station Option A: two MikroTik CRS804 switches, four Sparks each plus one Station port each, joined by one 400G trunk that is 2:1 oversubscribed. Option B: one NADDOD N9200-32DC, all eight Sparks and both Station ports on one ASIC, 26 ports spare. Option C (status quo): one CRS804 full of Sparks, Station off the fast fabric. Option A · two CRS804s, one 400G trunk ~$2,600 list · both switches 4/4 ports used · trunk carries up to 800G of demand at 400G CRS804 #1 4× 400G QSFP56-DD · Marvell 98DX7335 CRS804 #2 4× 400G QSFP56-DD · Marvell 98DX7335 400G trunk · 2:1 +~1 µs · shared by 4 Sparks each side Spark 1Spark 2 Spark 3Spark 4 Spark 5Spark 6 Spark 7Spark 8 GB300 Station CX-8 · 2× 400G, one per switch Blue: 200G Spark lane (Q2Q56 breakout). Green: 400G Station port. Red: the one link every cross-half packet must share. Option B · one NADDOD N9200-32DC $10,459 list · 1 month lead · 6 of 32 ports used, 26 spare · every host one ASIC hop from every other N9200-32DC · 32× 400G QSFP-DD/QSFP112 · 12.8 Tbps · Enterprise SONiC 4 cages broken out 2×200G for the Sparks · 2 cages native 400G for the Station · 26 cages empty Spark 1Spark 2Spark 3Spark 4 Spark 5Spark 6Spark 7Spark 8 GB300 Station CX-8 · 2× 400G, both in Option C · status quo: one CRS804, Station on 10 GbE $0 more · all 8 Sparks on one hop · Station cannot join a multi-box model at a useful rate CRS804 · 8× 200G, all used GB300 Station · 10 GbE island

Quantifying "unavoidable inter-switch latency"

The vendor's sentence is correct. Here is what it costs. Nobody publishes a measured port-to-port latency for the CRS804, so the switch-hop figures are ASIC-class estimates; the NIC-to-NIC numbers are measured by other Spark owners.

One-way small-message latency by topology Three horizontal bars showing the estimated one-way RDMA small-write latency for a direct cable, one switch hop, and two switch hops, followed by a callout computing the per-token cost of the extra hop for tensor-parallel decode across eight Sparks. One-way latency, 2-byte RDMA write, Spark → Spark Scale: 120 px per µs. Solid = measured by other Spark owners (ib_write_lat). Hatched range = estimate. 01 µs2 µs3 µs4 µs5 µs Direct cable 1.5–2.0 µs measured (NIC + NIC + firmware) One CRS804 ≈2.5–3 µs: + cut-through ASIC ~0.5–1 µs + FEC ~0.15 µs Two CRS804s ≈3.5–4.5 µs (+0.7–1.5) Per token, worst case (TP=8 decode): ~2 all-reduces/layer × ~60 layers × ~1 µs extra hop ≈ 0.12 ms against a 30–50 ms token step → 0.25–0.4%. Pipeline parallel crosses the trunk once per token → effectively 0. NCCL's own small-collective floor on Sparks is ~40 µs per call, which dwarfs the hop either way.

Working through it:

Latency is not the reason to reject two switches. If Option A is rejected, it should be for the trunk, the ports, and the cabling, not the microsecond.

The actual Option A problem: the trunk

Four Sparks per side at 200G is 800G of potential offered load per direction across a single 400G link: 2:1 oversubscribed. A ring all-reduce can be laid out so only one ring edge crosses the trunk in each direction, which is 200G on a 400G link and fine. NCCL's tree and mixed algorithms are not that polite, and neither is any Station-to-cluster fan-out (weights, KV pages, expert tensors) that hits multiple Sparks on the far side at once.

The second problem is that both switches end up at 4/4 cages: 2 breakout cages for Sparks, 1 for the Station, 1 for the trunk. There is no room to add a second trunk link, a ninth Spark, or the Station's second port on the same switch. Every future change is a re-cable. The single-switch layout leaves 26 empty 400G cages.

The third is operational: two RouterOS switches means PFC/ECN lossless config kept in sync across two boxes and a trunk that also has to carry PFC correctly. Doable; not free.

Price, power, noise, gotchas

Figures with a source are marked; the rest are labelled estimates. CRS804 list price and noise are from ServeTheHome's review; power from the MikroTik manual; N9200 figures from NADDOD's product page and quote.

ItemOption A · 2× CRS804-4DDQ-hRMOption B · 1× N9200-32DCOption C · 1× CRS804 (status quo)
Switch cost (list)~$2,590 ($1,295 each; street ~$1,100)$10,459 (quoted)~$1,295, already ordered
Extra cabling1× 400G QSFP-DD DAC for the trunk (~$89) + 1× Q112 DAC for the second Station portReuses the 4 Q2Q56 breakouts and the Q112 DAC; second Station DACnone
AvailabilityIn stockNext batch late October 2026; "1 month" on the quote—
ASICMarvell 98DX7335 (Prestera 7K, 1.6 Tbps; designed for 5G fronthaul/edge)Not stated on the product page. Sibling N9200-64DC is Broadcom Tomahawk 4Marvell 98DX7335
Switching capacity2× 1.6 Tbps, 400G between them12.8 Tbps, single hop1.6 Tbps
Ports used / total8/8 cages (zero spare)6/32 cages (26 spare)4/4 (Station excluded)
Cross-fabric latency+~1 µs for 16 of 28 Spark pairs est.One ASIC hop for everythingOne hop; Station not on fabric
Power, idle~29 W each measured → ~60 WNot published. Class estimate for a 32×400G 1U: several hundred watts est.~29 W measured
Power, max rated123 W each (92 W without optics)Not published; 10 A @ 200–240 VAC supplies123 W
AC input100–240 VAC, dual hot-swap200–240 VAC only (or 240 VDC), dual hot-swap100–240 VAC
NoiseLow-40s dBA empty, 50+ with hot optics (measured, per unit). Two of them.Not published. 6 hot-swap fans, connector-to-power airflow, 40 °C rating: datacenter-class est.Low-40s dBA
DepthHalf-width desktop/rack, shallow509 mm (20 in) deep 1UShallow
OSRouterOS, known PFC/ECN recipe (STH ran 8× GB10 NCCL on it)Enterprise SONiC by NADDOD; standard RoCE toolchain, new box, new to usRouterOS
Management16× 1GbE + consoleConsole + RJ45 mgmt + one 25G SFP2816× 1GbE
The 240 V gotcha. The N9200-32DC lists AC input as 200–240 VAC ~10 A. There is no 120 V mode. That means a dedicated 240 V circuit and a PDU or receptacle to match, in a rack that today is on 120 V with the Station's own 20 A feed. It is a solvable electrical job, but it is a job, and it belongs in the price. The CRS804 runs on anything from 100 to 240 V.
The noise gotcha. The CRS804 is a quiet box by 400G standards and the review numbers back that up. The N9200 is a 1U datacenter switch with six fans and a 40 °C intake rating; NADDOD does not publish dBA, and the sibling N9200-64DC datasheet lists a 2,400 W maximum. Assume it will need to live somewhere you do not sit, or behind a door. If the rack shares a room with people or dogs, that is a real cost.

What the vendor got right and wrong

What happens next

Nothing has been cancelled or ordered. The open questions, in order:

  1. Is there a 200G-class switch with 8× QSFP56 plus 2× 400G that is 120 V-capable and quieter than a Tomahawk 1U? NADDOD's own N8600-24QC8DC is $11,999 and 100–240 V, which is more money for less fabric. Other candidates need checking.
  2. If the answer is the N9200, what does a 240 V 20 A circuit at the rack cost, and where does the box physically live?
  3. If the answer is two CRS804s, lay out the ring so that only one edge per direction crosses the trunk, put one Station port on each switch, and accept that every future port change is a re-cable.

Sources used for this note: NADDOD product page #106164 and the September 30 quote email; ServeTheHome's CRS804-4DDQ-hRM review (price, idle power, noise, GB10 NCCL result); MikroTik hardware manual for the CRS804 (PSU input, 92/123 W); Marvell Prestera 98DX73xx product brief; three independent DGX Spark ib_write_lat write-ups (1.45–2.01 µs); Broadcom Tomahawk 4 documentation (~450 ns L3 latency, +~150 ns FEC) as the class reference for the "one ASIC hop" bar. Per-token arithmetic is an estimate and labelled as such.