← All Topics

M3 Ultra 512 Benchmarks

Public notes from one Mac Studio M3 Ultra with 512 GB unified memory. Recipes, dated measurements, and the configs that lost. This is the Apple Silicon 512 GB page to share.

If you only read one: Ornith 1.5 vs GLM-5.2 on the M3 Ultra — same-box A/B, same tools. Ornith is ~2× decode and finished 32k prefill. GLM-5.3 weights were not out when that ran.

GB300 and Spark (GB10) posts are different machines. Those live under DGX Station GB300 and Local LLMs.

What this Studio has actually served

WorkloadWhat is provenRead
Ornith 1.5-397B Q6_K 246,435-token cold prefill in 33.8 min at 122 tok/s. Decode 29.7 tok/s. Native 262k serve. Ornith 1.5-397B
Ornith 1.5 vs GLM-5.2 Sequential A/B, same tools and agent loop. Ornith ~2× decode and finished 32k prefill. GLM-5.3 weights not out. same-box A/B
oMLX PR #3059 DeepSeek-V4 ANE prefill Gains replicate after a one-line scheduler fix. Round 1: 625.3 vs 581.9 tok/s prefill. Round 4 tuner: oQ2.5e +8.8%/+8.2% PP, fp8 +4.2%/+4.4%. Round 4 · original field test
oMLX dual-ANE prefill (stock 0.6.1) Ultra 8k: 451 GPU vs 532 ANE (+18%). Ultra 16k: 438 vs 520 (+19%). M5 fused mode lost on that box. three-Mac note
MiniMax H3 FL2VA 8-bit Works with mlx-serve PR #122 on this 512 GB Studio. Stock v26.8.1 dies on load. H3 on MLX
GLM-5.2 MXFP4 / Terminus-2 ~368 GB on disk, 76 shards. Terminus-2: 49% (37/76) at 16K, 4-bit llama.cpp. MLX soloheaven was dead for that bench. Later GGUF A/B picked Ornith. Terminus-2 · MLX recipe
GLM-5.1 465 GB model. DQ4plus-q8 OOM. BAAI official quant (381 GB, 130 GB headroom) was the stable fit. Packing an Elephant
oMLX DeepSeek V4 Flash 4-bit Tool calls first try. Prefix cache 3.4×. oMLX DSv4-Flash
Speculative decode on 397B 4-bit Long reasoning +58%. Large-context tool calls −88%. 4B draft study
Ornith 1.0-397B bartowski Q4_K_M GGUF ~213 GiB RSS and ~31 tok/s. MLX 8-bit 250K canary: 252,352 prompt tokens in 1,723.9 s. GGUF and MLX recipes
Kimi K2.6 DQ3 Negative result: this quant cannot split reasoning vs answer via a closing tag. The Closing Tag That Never Comes

Numbers are dated lab measurements from the linked posts, not leaderboard claims. Different engines, quants, harnesses, and context lengths are not automatically comparable.

Posts

Related