Someone just got a pair of DGX Station GB300s in and texted that the machines do not see the 10G cables or the QSFP cables. Both machines, both cable types. That is frustrating, and it is usually the rear panel, not two dead NICs.
Ubuntu owns two networks on the back of this box. A third jack looks like Ethernet and is not Ubuntu's. The QSFP cages are 400-gigabit ConnectX-8 ports, not a faster version of the 10G jack, and not the same cages as a DGX Spark.
Rear-port roles from NVIDIA's DGX Station software spec and the published Station rear-panel list. Interface names are from our MSI box, not a promise about every OEM image.
Split the symptom before you buy another cable
"Doesn't see the cable" is three different failures. Run this on the Station, not from memory of the switch UI.
ip -br link
lspci -nn | grep -iE 'mellanox|aquantia|atlantic'
sudo ethtool IFACE
Replace IFACE with whatever ip -br link actually printed. On our MSI image the host 10G came up as eno3 and the ConnectX-8 pair as enP1p3s0f*. Yours may differ. Identify the chip, not the name I used.
| What you see | What it means | Do this |
|---|---|---|
No Mellanox and no Aquantia/Atlantic in lspci | The NIC is not enumerated. A new cable will not fix that. | Stop. Do not open the chassis. Save lspci -nn and dmesg, and call the OEM. |
Device is there, ethtool says link detected no | The port exists. The cable, the cage, or the far-end speed does not match. | Reseat until the latch clicks. Try the other cage. Read the cable part number. |
| Link is up, speed is 1000Mb/s | You are on the BMC jack, or the switch port is 1G. | Move the host cable to the 10GbE RJ45. Confirm the switch port is 10G. |
ConnectX-8 is in lspci, interfaces exist, carrier down, no 400G peer | Normal. Our CX8 pair sits down until a matching cable is attached. | Do not RMA a dark QSFP cage that still enumerates. |
The 10G jack
NVIDIA's station software spec is one 10GbE RJ45 for in-band management, plus a separate 1GbE RJ45 that belongs to the BMC. The published rear-panel list names the host chip as an AQC-113C and the second RJ45 as the BMC. Two jacks, two owners.
If the host 10G "isn't seen," the usual miss is the BMC jack. It can light up at 1G for the management web UI and never appear as a host interface. Plug the host cable into the other RJ45, then ask ethtool for Speed: 10000Mb/s and Link detected: yes. A link at 1000Mb/s is a real link. It is not 10G.
If the Aquantia device is in lspci and carrier is still down, reseat the RJ45, try a known-good Cat6, and check that the switch port is actually a 10G port. A 1G switch port will link at 1G and feel like the Station refused the cable.
The QSFP cages
Those two cages are a ConnectX-8, two ports of 400GbE, QSFP112. They are not the 10G network. They are not a DGX Spark ConnectX-7, which is a pair of 200G ports and a different PHY.
A 10G SFP+ cable, a cheap QSFP-to-SFP+ adapter, a 100G QSFP28 DAC, or a Spark QSFP56 DAC will sit dark in this cage. I have not verified a 10G QSA on our Station, and I would not start there. Spark owners have a separate writeup where a Mellanox-compatible QSA plus forced 10G and autoneg off brought up a ConnectX-7 on a UniFi switch. That is a Spark note. Copying it onto a Station CX8 is how you lose an afternoon.
What I would do instead:
- Confirm ConnectX-8 is in
lspciwith no cable installed. Both cages should enumerate even when they are dark. - If you want them to link, use a QSFP112 DAC or AOC rated for 400G, or a module the CX8 firmware actually supports. Direct Station-to-Station is the simple case: matching cable, MTU 9000, static addresses, no default route on that rail. Leave the 10GbE as the management path.
- The far end has to be the same speed. A 10G or 25G switch port will not light a 400G cage.
- Reseat until the latch clicks, then try the other cage. QSFP112 cages are stiff. "I pushed it" is not seated.
If lspci shows no Mellanox at all, stop. That is not a cable, and it is not a reason to pull the side panel on a machine that is still under warranty. Spark has an internal-cable and hotplug failure mode on its ConnectX-7. I am not going to tell you that is what a Station is doing.
Power, so the rest of the week is boring
The rear IEC rocker is standby, not the OS boot. The front button can sit dead for 10 to 15 seconds while the BMC wakes. Fans and the pump come up after that. That delay was normal on our box. If the front button stays dead and the PSU LED and pump stay dark after a firmly seated C19 on a dedicated 20A circuit, stop. Do not hold the button, and do not open the chassis.
Every cord and strip in the path to the PSU should be 20A-rated. A consumer surge strip is for the monitor and the keyboard, not a 1600W supply.
On the PSU we have, the 1600W band is 115-240V. The 1300W band is 100-114V. US 120V is already in the 1600W band. I would not convert a 20A circuit to 240V to chase watts. It changes current and a bit of efficiency. It does not unlock a hidden 300W.
Leave the little display GPU alone. Ours is an RTX PRO 4000 SFF, 70W, and it is how the desktop gets a picture. A 6000-class card in that slot eats the GB300 power budget. I would not do that swap for "extra inference." The next dollar is NVMe, or a second Station on a real 400G cable.
Picture, and the two logins
The desktop picture comes from the add-in GPU. On the 4000 SFF that is four Mini DisplayPort connectors, not HDMI. Mini-DP looks like mini-HDMI. Check for the DP logo. The chassis USB-C port is data. The BMC mini-DP is a low-resolution management console, not the OS. NVIDIA's own software spec says the BMC mini-DP is not the primary desktop display.
If the monitor is HDMI, use an active Mini DisplayPort to HDMI adapter. Passive dongles often fail the handshake. Do not hang a JetKVM or an HDMI capture box on the BMC mini-DP.
There are two passwords, and they are not interchangeable. The Ubuntu account is the one on the vendor's system sheet. It is not in NVIDIA's public docs, and I am not going to guess it here. The first-boot dialog that wants to run bmc-user-setup creates a BMC account, not an Ubuntu account. Cancel is safe. Finish it later with sudo /usr/libexec/bmc-user-setup so the factory BMC login does not stay on the LAN. If the dialog rejects a password that just logged you in, check the symbol, then wait out a lockout. Three bad tries locks auth for about 10 minutes, and the correct password fails the same way during that window. Prove it with sudo -v in a terminal instead of the dialog.
Put the BMC on a trusted network. Do not put it on the internet.
nvidia-smi cannot see the GPU
We hit this on first boot, after the image came up on a kernel that did not have its matching NVIDIA module package. nvidia-smi said it could not talk to the driver. The fix was the module package for the running kernel, then a reboot:
sudo apt-get install -y linux-modules-nvidia-595-open-$(uname -r)
sudo reboot
After any later apt upgrade, check that the package list still contains the running kernel before you reboot. A new kernel without that module brings the same dead GPU back. Skip Ubuntu Livepatch on the nvidia-64k kernel. Let the DGX OS repo own the driver. Hold a big upgrade until nvidia-smi is healthy.
When it works, GPU 0 on our box is the display card. GPU 1 is the GB300. Containers that should use the big GPU need to be pointed at it. The 4000 is not an inference device.
A container is not a wall on this kernel
DGX OS is Ubuntu 24.04 with NVIDIA's kernel. Fetched October 2, 2026, Ubuntu's tracker still marks that kernel vulnerable to CVE-2026-80521. It is a use-after-free in the kernel's local-socket garbage collector. Code already running inside an ordinary container can become root on the host, because the calls it needs are ones Docker allows by default. It is not a remote bug. Nothing on the internet reaches it unless something is already executing on the box.
The packages a Station actually runs are still listed Vulnerable on 24.04: linux-nvidia, linux-nvidia-6.17, and linux-nvidia-7.0. Our box has been on both the 6.17 and 7.0 nvidia-64k lines. Check uname -r, then the tracker. An apt upgrade will not close this until Ubuntu marks the package fixed. Do not cherry-pick the upstream patch. Ubuntu's page says not to.
Until that row says fixed:
- Do not run a container you do not trust. That includes a model repo passed to
--trust-remote-codethat you did not pin, and a:latestimage you pull because a thread said to. - Do not publish Docker ports on every interface. Docker installs firewall rules ahead of ufw, so a published port is reachable from anything that can reach the machine, whatever ufw says.
- Keep the BMC off the internet. Same rule as above.
What this looked like in our own lab is in the September 23 security note.
The memory number you will misread
NVIDIA's software spec lists 496 GB of LPDDR5X for the Grace side. It does not, on that page, tell you how much HBM nvidia-smi will print. A March 2026 deskside roundup lists the workstation GPU as 252 GB HBM3e, about 12% under a full server GB300, and the CPU as 496 GB.
On our box nvidia-smi reports 256,703 MiB. That is 250.7 GiB. It is not 256 GB, it is not 269 GB, and it is not 288 GB of usable HBM. 269 GB is what you get if you convert that same MiB figure to decimal gigabytes. 288 GB is the server-part spec people quote. Do not RMA the card because the tool did not print 288.
Our host reports about 494 GiB and no swap. If MemTotal looks enormous and then drops after a driver setting, that can be coherent-memory mode, not missing RAM. I would not change that setting in the first week.
Disks
The OS pair is a software RAID 1. Do not add a drive to that array, and do not identify disks by nvme0 or nvme1. Those names reshuffle when you add drives. Use the serial in /dev/disk/by-id. On our box, touching the wrong member would have wiped the OS.
The root volume is about 1.9 TB and it fills fast once you stage a couple of big checkpoints. Stage weights on local NVMe before you serve them. A network mount will make the first load look hung, because the loader page-faults the file through the share.
What to run first
Prove the GPU with a model that fits before you try one bigger than HBM. When you do go bigger, the path that worked for us was vLLM reading Grace in place (UVA offload), not a stack that copies a whole expert layer across the link for every token. That second path capped near a few tokens per second on this box and did not get better by tuning.
The first boot of a big vLLM image can sit at high GPU with no new log line while FlashInfer autotunes. That is not automatically a hang, and high GPU is also not proof that it is making progress. If the container has exited, or the host is out of memory, it is not autotune. Kill it then. Otherwise let the first shape finish, and keep the cache it writes. The next boot of the same shape should be much shorter.
What I would send back
If the cables are still dark after the jack check, I want four lines, not a theory: ip -br link, the lspci grep, ethtool on the host 10G interface, and the part number printed on the QSFP cable. No carrier plus no Mellanox device is an OEM call. No carrier plus a ConnectX-8 in lspci is a cable or a switch-speed problem. A link at 1G means the host cable is in the wrong jack.
Sources
- NVIDIA DGX Station software spec: one 10GbE RJ45, two 400GbE QSFP ports on a ConnectX-8, a separate 1GbE BMC, desktop video from the add-in GPU, BMC mini-DP is not the OS display, 496 GB LPDDR5X, OS drives in RAID 1.
- ServeTheHome deskside Station roundup, March 20, 2026: rear ports listed as 2x 400GbE QSFP112, 10GbE RJ45 on an AQC-113C, 1GbE BMC, 1600W PSU, workstation GPU listed as 252 GB HBM3e.
- NVIDIA developer forum, DGX Spark / ConnectX-7: a 10G QSA thread. Different NIC, different cage. Cited so you do not treat that fix as a Station procedure.
- Ubuntu CVE-2026-80521, package status read October 2, 2026: 24.04
linux-nvidia,linux-nvidia-6.17, andlinux-nvidia-7.0still Vulnerable. Not a remote bug. No exploit steps on this page.
Our interface names, the 256,703 MiB reading, the 4000 SFF, and the kernel-module miss are from this lab's MSI Station, not from the spec sheet.
Related: what our Station is actually running, and the GB300 topic index.