GB300 Links
A reading list. The web is finally publishing deskside GB300 coverage that is not a partner datasheet. Two pieces landed in front of us today: Ahmad Osman sitting next to a DGX Station on MTS, and StorageReview's full MSI XpertStation WS300 review. Everything below is their numbers, labeled. Ours live on the GB300 topic page.
Desk is not rack. DGX Station is one Grace + one Blackwell Ultra, published as 252 GB HBM3e + 496 GB LPDDR5X = 748 GB coherent. GB300 NVL72 is a 72-GPU liquid rack. Aggregator pages that list “GB300, 288 GB HBM” are talking about the rack chip, not the box under a desk. ServeTheHome's Computex look at the Super AI Station is the cleanest public statement of why: Station SKUs enable 7 of 8 HBM3e stacks.
Start here
StorageReview WS300 is the independent review. MTS + Ahmad Osman at 18:19 is the Station in a room.
Spec sheet
NVIDIA's own Station page: 252 GB HBM3e at 7.1 TB/s, 496 GB LPDDR5X at 396 GB/s, 900 GB/s C2C, 1,600 W, ConnectX-8 up to 800 Gb/s.
The 748 GB trick
Coherent address space, not 748 GB of HBM. Weights that miss HBM still cross C2C and pay Grace's ~396 GB/s. StorageReview measured that path.
Our lab
Recipes and dated measurements: DGX Station GB300. Different engines and quants. Do not paste their tok/s next to ours.
The two that started this page
Ahmad Osman on MTS, with a Station in the studio
Why Frontier AI Labs Are Terrified of Open-Source — MTS, September 12, 2026, 39:15. Guest is Ahmad Osman (Osmant). They physically brought a DGX Station into the studio and then could not plug it in: 20 A circuit, too hot for the set. Jump to 18:19 for the hardware block.
Auto-captions, not a signed bench. Osman describes the box as a GB300 / B300 machine: 252 GB HBM3e at 7.1 TB/s plus 496 GB LPDDR5X at ~396 GB/s. He says he is running GLM-5.3 Flash and Qwen 3.8 Next Flash, and claims on-box serving on the order of thousands of tokens per second with 10–12 concurrent requests at about 90–100 tok/s per request. Treat that as a practitioner talking on camera, not a harness. The rest of the interview is the sovereignty argument: you need enough local inference to notice when a cloud API has been quantized out from under you.
Transcript via YouTube auto-captions (English, generated). Timestamp 18:19 is the hardware turn James sent.
StorageReview: MSI XpertStation WS300
MSI XpertStation WS300 Review: 748GB of Coherent Memory and 20 PetaFLOPS on a Desk — Divyansh Jain, August 24, 2026. This is the review to send someone. NVIDIA supplies the Grace Blackwell Ultra baseboard; MSI supplies chassis, loop, PSU, storage layout, I/O, and support. They tested remotely on an MSI-hosted unit with no RTX PRO installed, so the Superchip had the full accelerator budget.
Two caveats they printed themselves, which most recaps drop:
- Speculative-decode runs set forced acceptance equal to the draft length. Those tok/s numbers are a best-case ceiling, not live serving.
- 748 GB is one coherent address space. HBM stays the hot tier. Grace is the overflow tier. Crossing C2C is a real cost, and they measured it.
| What they measured | Their number | Why it matters |
|---|---|---|
| MAMF, dense NVFP4 / FP8 / BF16 | 6,134 / 3,909 / 1,967 TFLOPS | 4–5× a 600 W RTX PRO 6000; 17–19× a DGX Spark, stable across precisions. |
| HBM device-local read | 6,875 GB/s (~6.9 TB/s) | Within a few percent of the 7.1 TB/s rating. |
| Grace → HBM / HBM → Grace | 390 / 382 GB/s | C2C is rated 900 GB/s; the limiter they measured is Grace LPDDR at 396 GB/s. |
| Thermals under their load | GPU 71 C, CPU 65 C, GPU 1,292 W | MSI rates the loop 1,400 W across CPU+GPU. One 1,600 W PSU, C19, dedicated 20 A circuit. |
| DeepSeek V4 Flash (0731), native FP8, 512/512 | 149 tok/s C1; 1,766 tok/s C32 | Fits in HBM. Their earlier teaser tweet used the C32 figure. |
| MiniMax M2.7 NVFP4, 512/512 | 193 tok/s C1; 4,801 tok/s C128 | 125 GB in HBM. Comfortable fit. |
| MiniMax M3 NVFP4 | 223 GB HBM + 21 GB experts on Grace | “Barely fits.” Prefill-heavy falls off after C2. |
| GLM-5.2 NVFP4 (433 GB ckpt) | 218 GB HBM + 216 GB on Grace; 36→139 tok/s C1→C32 | Does not fit. Both prompt lengths track, so they call expert migration the limiter. Sweep dies at C32: no HBM left for more KV. |
| Nemotron-3-Ultra 550B NVFP4 | 114 GB on Grace; 43→168 tok/s C1→C32 | Largest model they loaded. Same Grace-bound pattern. |
| GPT-OSS-20B vs RTX PRO 6000 vs Spark, 512/512 @ C128 | 22,161 / 9,000 / 1,469 output tok/s | Shared-model comparison: Station 2.3–3× the card and 14–19× Spark at full concurrency; gap opens toward 6× the card on long prompts. |
They also note the optional RTX PRO is not free performance on top. vsloshd gives the discrete card priority; when both are busy, GB300 clocks come down. Mini DisplayPort is BMC video at 1024×768. If you want a real display, you add an RTX PRO. That is the same story ServeTheHome told from the other side of the plexiglass.
Conclusion, in their words: one or two of these next to a desk, as a development node, against cluster-reservation cost. Four of them and you should be shopping an eight-way B300 server instead.
Official
- NVIDIA DGX Station product page — 748 GB coherent, “up to 20 petaFLOPS,” models “up to 1 trillion parameters,” ConnectX-8, link two Stations. Specs table: 252 GB HBM3e | 7.1 TB/s; 496 GB LPDDR5X | 396 GB/s; NVLink-C2C 900 GB/s; FP4 Tensor Core 20 | 15 PFLOPS (sparse | dense); 1,600 W; 2× QSFP112 400 Gb/s; 4× M.2 Gen 5; MIG 7. There is now a “DGX Station for Windows” announcement on the same page.
- DGX Station Development Guide — System Overview — the short architecture note: 72-core Grace Neoverse V2, 114 MB L3, GB300 GPU “up to 252 GB HBM3e,” Ubuntu 24.04 with NVIDIA AI Developer Tools.
- DGX Station datasheet (PDF) — same numbers in vendor form.
- GTC 2025 announcement — the original Station reveal. Early copy said 784 GB coherent; shipping pages say 748 GB. If you still see 784, it is leftover launch language.
- GB300 NVL72 and DGX GB300 — the rack, not the desk. 72 Blackwell Ultra GPUs, 36 Grace CPUs, 20 TB GPU memory. Keep these tabs closed when you are shopping a Station.
OEM deskside boxes
NVIDIA sells the Superchip and the software stack. Partners sell the tower. Internals rhyme. Differentiation is cooling, storage population, display-card option, rack-kit, and who answers the phone.
| OEM | What the public web currently says | Read |
|---|---|---|
| MSI XpertStation WS300 | VideoCardz (August 31, 2026): shipping through ASI, D&H, and Newegg; Newegg list $99,999, marked out of stock when they wrote it. Guru3D had $99,900 two days earlier. StorageReview still had no published MSI list price in the August 24 review. | VideoCardz |
| ASUS ExpertCenter Pro ET900N G3 | VideoCardz (June 15, 2026): same 748 GB split, ConnectX-8, “future Windows support.” UK list £117,599.99. An earlier ASUS listing used 784 GB; the updated spec is 748. | VideoCardz |
| HP ZGX Fury AI Station | Notebookcheck (September 10, 2026): same GB300 Superchip, 496 GB LPDDR5X + 252 GB HBM3e, Red Hat collaboration, optional discrete RTX PRO in the x16 slot. They put the system “somewhere in the range of $100,000.” | Notebookcheck |
| Dell Pro Max with GB300 | TechRadar sponsored comparison (September 10, 2026): marketing numbers — 120B to 1T parameters, “up to 150 concurrent agents,” MaxCool liquid cooling, a Signal65/Futurum TCO claim versus cloud APIs. Not a hardware review. | TechRadar |
| Supermicro Super AI Station | ServeTheHome (June 27, 2026): plexiglass Computex unit. ~$125k Newegg affiliate tag on DGX Station as a class. 7-of-8 HBM stacks, 252 GB / 7.1 TB/s, four SOCAMMs, GB300 is not graphics-capable, optional RTX PRO for a real display, 5U rack kit as a differentiator, 1.6 kW not everyone wants at a desk. | ServeTheHome |
| Exxact (MSI chassis) | That is the box in this lab. Hardware notes we actually ran: filling the empty CX8 M.2s, day one. | Ours |
Other coverage worth keeping
- ServeTheHome, MSI GTC 2026 booth — open-chassis photos. Linked from our M.2 post because MSI still does not publish a hardware install manual, only a BIOS PDF.
- LinuxGizmos WS300 internals — labeled internals and the user-serviceable vs not list. The origin 520'd while this page was being written; keep the URL, do not treat a Cloudflare error as a takedown.
- StorageReview Computex short — CoolIT-looking loop, copper cold plates, MSI's slot-load RTX PRO 6000 platform with QD fittings. One minute, useful before you open a side panel.
- StorageReview's own Reddit recap — same review, compressed. Useful if you only want the HBM-boundary experiment in one paragraph.
- Guru3D, August 29, 2026 — $99,900 retail note, “two units supported,” not a bench.
- MSI WS300T60L product page and BIOS user guide PDF — the official paper. No hardware install section.
Skip, or read with a hard filter. Cloud GB300 price aggregators (288 GB HBM, NVL72, $/GPU-hr) are a different product. Vendor TCO pages that convert a Station into “150 concurrent agents” or “87% lower token spend” are not measurements. Recap sites that reprint NVIDIA's 20 PFLOPS without saying sparse FP4 are doing brochure math. NationalPC-style pages that rate the WS300 9.2/10 from the spec sheet are not reviews.
How to read any of this against our lab
StorageReview's GLM-5.2 offload numbers and Osman's on-camera GLM-5.3 Flash claims are not substitutes for a recipe. Different checkpoints, quants, engines, acceptance cheats, and concurrency. If you want what one Station in this house actually served, with the configs that lost: DGX Station GB300.
The useful public fact, independent of whose tok/s you trust: the deskside SKU is a 252 GB HBM machine with a large, slower Grace tier attached by C2C. Everything interesting about serving on it is which side of that 252 GB line you are on.
Related here
- DGX Station GB300 — recipes and dated measurements from this lab.
- Filling the empty M.2s — CX8 slots, RAID0
/models, nevermdadmbynvmeN. - Day one — arrival, kernel trap, first token.
- Local LLM stack — where the Station sits relative to the Spark pair.