Switching 8 DGX Sparks and a GB300 Station: CRS804 ×2 vs one 32×400G switch
CreatedLast updated
Decision note · James Meadlock & Milo (James's AI agent) · written on claude-fable-5-1 (Anthropic), default reasoning · prompted by a vendor reply from NADDOD
Where this stands: the fabric is eight DGX Sparks at 200 GbE each plus one GB300 Station with two 400 GbE ports. One MikroTik CRS804 has exactly eight 200G lanes and nothing left for the Station. NADDOD's answer was "buy our N9200-32DC" and "two CRS804s add unavoidable inter-switch latency." Both statements are true. The latency is about 1 µs per extra hop, which is under 0.5% of a decode step. The real reasons to prefer one switch are trunk bandwidth, zero spare ports, and the fact that the Station wants two uplinks. The real reasons to hesitate are $10,459, a late-October batch, six datacenter fans, and 200–240 VAC input only. No order has changed yet.
~1 µs
added one-way latency per extra switch hop estimate
<0.5%
of a 30–50 ms decode step, worst-case TP=8 estimate
Each DGX Spark exposes one ConnectX-7 port at 200 GbE. The GB300 Station has a ConnectX-8 with two QSFP112 cages at 400 GbE, running in Ethernet mode. The CRS804-4DDQ-hRM has four QSFP56-DD cages at 400G; each one breaks out to 2×200G with a Q2Q56 splitter DAC. Four cages × 2 = eight 200G lanes. Eight Sparks. Full. The Station has nowhere to plug in.
Until recently that was fine. The Sparks talked to each other and the Station was a separate island on 10 GbE. The plan changed once it became clear that a single model can usefully span HBM on the Station, Grace memory on the Station, and the Spark cluster together — a Kimi K3 NVFP4 style placement — and at that point the Station needs to be on the same fabric at a rate that is not 10 GbE.
The three options, drawn
Quantifying "unavoidable inter-switch latency"
The vendor's sentence is correct. Here is what it costs. Nobody publishes a measured port-to-port latency for the CRS804, so the switch-hop figures are ASIC-class estimates; the NIC-to-NIC numbers are measured by other Spark owners.
Working through it:
Floor with no switch at all: 1.45–2.01 µs typical for a 2-byte ib_write_lat between two directly cabled Sparks (three independent write-ups, all ConnectX-7 RoCE). That is two NICs and their firmware.
One CRS804 in the path: the Marvell 98DX7335 is a cut-through ASIC; the vendor community pegs it at sub-microsecond cut-through and 2–5 µs store-and-forward for full frames. Add ~150 ns for the second link's PAM4 FEC decode and 5 ns/m of cable. Call it +0.7–1.0 µs, so ~2.5–3 µs one-way.
Two CRS804s: one more ASIC traversal plus one more FEC'd link, +0.7–1.5 µs. Roughly 3.5–4.5 µs one-way. That is a 25–50% increase on the raw wire number, which sounds dramatic until you see what it is a fraction of.
Only cross-half pairs pay it. With 4 Sparks per switch, 16 of the 28 Spark pairs straddle the trunk; 12 do not. The Station can be one hop from everything if it puts one 400G port in each switch, at the price of two subnets or L3 to steer traffic to the right port.
What it does to a token: tensor-parallel decode across all eight is the worst case, ~2 collectives per layer × ~60 layers ≈ 120 small all-reduces per token. At ~1 µs each that is ~0.12 ms per token against 30–50 ms token steps: under half a percent. Pipeline parallel, which is what most Spark clusters actually run because 200G is bandwidth-starved for TP, crosses the trunk exactly once per micro-batch. And the measured NCCL small-message floor on Sparks is around 40 µs per collective, so the switch hop is buried under software regardless.
Latency is not the reason to reject two switches. If Option A is rejected, it should be for the trunk, the ports, and the cabling, not the microsecond.
The actual Option A problem: the trunk
Four Sparks per side at 200G is 800G of potential offered load per direction across a single 400G link: 2:1 oversubscribed. A ring all-reduce can be laid out so only one ring edge crosses the trunk in each direction, which is 200G on a 400G link and fine. NCCL's tree and mixed algorithms are not that polite, and neither is any Station-to-cluster fan-out (weights, KV pages, expert tensors) that hits multiple Sparks on the far side at once.
The second problem is that both switches end up at 4/4 cages: 2 breakout cages for Sparks, 1 for the Station, 1 for the trunk. There is no room to add a second trunk link, a ninth Spark, or the Station's second port on the same switch. Every future change is a re-cable. The single-switch layout leaves 26 empty 400G cages.
The third is operational: two RouterOS switches means PFC/ECN lossless config kept in sync across two boxes and a trunk that also has to carry PFC correctly. Doable; not free.
Price, power, noise, gotchas
Figures with a source are marked; the rest are labelled estimates. CRS804 list price and noise are from ServeTheHome's review; power from the MikroTik manual; N9200 figures from NADDOD's product page and quote.
Item
Option A · 2× CRS804-4DDQ-hRM
Option B · 1× N9200-32DC
Option C · 1× CRS804 (status quo)
Switch cost (list)
~$2,590 ($1,295 each; street ~$1,100)
$10,459 (quoted)
~$1,295, already ordered
Extra cabling
1× 400G QSFP-DD DAC for the trunk (~$89) + 1× Q112 DAC for the second Station port
Reuses the 4 Q2Q56 breakouts and the Q112 DAC; second Station DAC
none
Availability
In stock
Next batch late October 2026; "1 month" on the quote
—
ASIC
Marvell 98DX7335 (Prestera 7K, 1.6 Tbps; designed for 5G fronthaul/edge)
Not stated on the product page. Sibling N9200-64DC is Broadcom Tomahawk 4
Marvell 98DX7335
Switching capacity
2× 1.6 Tbps, 400G between them
12.8 Tbps, single hop
1.6 Tbps
Ports used / total
8/8 cages (zero spare)
6/32 cages (26 spare)
4/4 (Station excluded)
Cross-fabric latency
+~1 µs for 16 of 28 Spark pairs est.
One ASIC hop for everything
One hop; Station not on fabric
Power, idle
~29 W each measured → ~60 W
Not published. Class estimate for a 32×400G 1U: several hundred watts est.
~29 W measured
Power, max rated
123 W each (92 W without optics)
Not published; 10 A @ 200–240 VAC supplies
123 W
AC input
100–240 VAC, dual hot-swap
200–240 VAC only (or 240 VDC), dual hot-swap
100–240 VAC
Noise
Low-40s dBA empty, 50+ with hot optics (measured, per unit). Two of them.
Not published. 6 hot-swap fans, connector-to-power airflow, 40 °C rating: datacenter-class est.
Low-40s dBA
Depth
Half-width desktop/rack, shallow
509 mm (20 in) deep 1U
Shallow
OS
RouterOS, known PFC/ECN recipe (STH ran 8× GB10 NCCL on it)
Enterprise SONiC by NADDOD; standard RoCE toolchain, new box, new to us
RouterOS
Management
16× 1GbE + console
Console + RJ45 mgmt + one 25G SFP28
16× 1GbE
The 240 V gotcha. The N9200-32DC lists AC input as 200–240 VAC ~10 A. There is no 120 V mode. That means a dedicated 240 V circuit and a PDU or receptacle to match, in a rack that today is on 120 V with the Station's own 20 A feed. It is a solvable electrical job, but it is a job, and it belongs in the price. The CRS804 runs on anything from 100 to 240 V.
The noise gotcha. The CRS804 is a quiet box by 400G standards and the review numbers back that up. The N9200 is a 1U datacenter switch with six fans and a 40 °C intake rating; NADDOD does not publish dBA, and the sibling N9200-64DC datasheet lists a 2,400 W maximum. Assume it will need to live somewhere you do not sit, or behind a door. If the rack shares a room with people or dogs, that is a real cost.
What the vendor got right and wrong
Right: one switch is the cleaner design; the Station ideally gets both 400G ports on the same fabric; a 32×400G box leaves room to grow without redesign; the existing NADDOD breakout DACs carry over.
Right but unquantified: "additional latency between the switches is unavoidable." About a microsecond, worth under half a percent of a decode step. True, and not the deciding factor.
Backwards: "the CRS804 is based on a traditional data center networking chip." The 98DX7335 is a Prestera 7K part marketed for 5G fronthaul aggregation and carrier edge, with TSN and PTP as its headline features. It is a worse RoCE pedigree than the sentence implies, not a different flavor of the same thing. It still works: ServeTheHome ran NCCL across eight GB10s on it once PFC was on.
Omitted: the ASIC in the N9200-32DC, its power draw, its noise, and that it needs 240 V.
Sales context to keep in view: the reply is from a vendor recommending its own new product, launched this month, with the next batch shipping late October. The recommendation is not wrong; it is just not disinterested.
What happens next
Nothing has been cancelled or ordered. The open questions, in order:
Is there a 200G-class switch with 8× QSFP56 plus 2× 400G that is 120 V-capable and quieter than a Tomahawk 1U? NADDOD's own N8600-24QC8DC is $11,999 and 100–240 V, which is more money for less fabric. Other candidates need checking.
If the answer is the N9200, what does a 240 V 20 A circuit at the rack cost, and where does the box physically live?
If the answer is two CRS804s, lay out the ring so that only one edge per direction crosses the trunk, put one Station port on each switch, and accept that every future port change is a re-cable.
Sources used for this note: NADDOD product page #106164 and the September 30 quote email; ServeTheHome's CRS804-4DDQ-hRM review (price, idle power, noise, GB10 NCCL result); MikroTik hardware manual for the CRS804 (PSU input, 92/123 W); Marvell Prestera 98DX73xx product brief; three independent DGX Spark ib_write_lat write-ups (1.45–2.01 µs); Broadcom Tomahawk 4 documentation (~450 ns L3 latency, +~150 ns FEC) as the class reference for the "one ASIC hop" bar. Per-token arithmetic is an estimate and labelled as such.