The M5 generation added a matrix unit to every GPU core. Apple calls them Neural Accelerators; the MLX source calls the kernel family NAX. Apple's own numbers say prompt processing (prefill) got roughly 4× faster than M4. Prefill is exactly the part of local inference that hurts for agents, because an agent turn is mostly reading context: our production traffic runs about 58 prompt tokens for every completion token.
So the hypothesis was tempting: new silicon, young software, a compute-bound workload. Surely there are ops MLX is not routing through the new hardware yet, and surely fixing that is low-hanging fruit.
Before touching a kernel we did two things: read the MLX source, and ask a second model (GPT-6 Astra) to tear the plan apart. The source reading showed MLX already has NAX paths for dense GEMM, quantized GEMM, mixture-of-experts gather-GEMM, attention prefill, and gated delta-net. Astra's review said the plan was optimising the wrong number: a higher share of time on NAX kernels is neither necessary nor sufficient for lower latency. Measure latency, use coverage as an explanation. That reframing is the reason this post has a clean conclusion instead of a spreadsheet of kernel names.
An LLM forward pass is mostly matrix multiplies. During prefill, thousands of tokens hit each weight matrix at once, so the multiply is big and square-ish: compute-bound. During decode, one token at a time hits each matrix: a skinny multiply that is bound by memory bandwidth, not math. The accelerators help the first case and do nothing for the second. That is why Apple quotes 4× for prefill and ~1.2× for generation.
The obvious tool is Metal frame capture: mx.metal.start_capture() writes a .gputrace that Xcode's GPU Frame Debugger can replay with per-kernel timings. We got that far. A 482-token prefill on a 4B model produced a 3.8 GB bundle of 1,053 files. But the bundle is not scriptable: dispatches reference pipeline objects by hash, kernel names live in a 190 MB embedded metallib of everything MLX compiled (not what it ran), and there are no timings at all until Xcode replays the capture. Capturing also slows the run 10× (100 ms to 1,060 ms), so captured runs are never timing data. Astra had flagged exactly this risk. It was right.
So we went one level up. The harness monkey-patches the four MLX ops that carry the compute (quantized_matmul, gather_qmm, matmul, fast.scaled_dot_product_attention, plus gated_delta_update for hybrid models), records the shapes of every call, times each call with a GPU synchronize, and predicts which kernel MLX dispatched by applying the gating rules copied from the MLX source (quantized.cpp, matmul.cpp, scaled_dot_product_attention.cpp, main at 7916d8b). Those predictions are labelled predicted until someone opens one capture in Xcode and checks them; we have not done that yet.
Two models already on the machine, chosen for the dense/MoE contrast: Qwen3.8-27B (dense, a mixed 4/5-bit MLX quant) and gemma-4-26b-a4b (mixture of experts, 4-bit). Workload: a 16,384-token synthetic agent prompt, cold, prefilled in mlx_lm's default 2,048-token chunks; then a 512-token delta appended to the warm cache, which is what an agent turn looks like after a tool call. Wall times below are uninstrumented; the breakdown comes from a second, instrumented pass.
| Model | Cold 16K prefill | tok/s | Warm 512 delta | Time on NAX kernels (predicted) |
|---|---|---|---|---|
| Qwen3.8-27B dense | 19.4 s | 845 | 828 ms | 98.5% |
| gemma-4-26b-a4b MoE | 4.3 s | 3,850 | 416 ms | 98.3% |
The amber slice is real and already known upstream: for very narrow outputs (N=32 or 48, the per-head gate projections in Qwen3.5-style models) MLX's split-K heuristic picks a kernel that has no NAX variant. Issues #3584 and #4198 cover it with microbenchmarks. On a real model at real prefill sizes it is 1.5% of the step. Fixing it is worth a PR; it is not worth a project.
If 92% of the time is in one kernel family, the question becomes whether that kernel is good, not whether it is used. The cleanest test: take the exact matrix shapes the 27B model hits during a 2,048-token step, run each one through the quantized kernel and through plain fp16 matmul, and compare. A quantized kernel has to dequantize on the fly, so it should be a little slower than fp16 per FLOP. If it were much slower, that would be the fruit. Issue #3584 reports exactly that pathology at small M (1.5–1.8×).
| Shape (N × K) | What it is | Share of 27B matmul time | 4-bit | fp16 | ratio | 4-bit TFLOPS |
|---|---|---|---|---|---|---|
| 17408 × 5120 | FFN up/gate, 4-bit | 37.5% | 6.48 ms | 5.92 ms | 1.09 | 56.3 |
| 5120 × 6144 | attention out, 5-bit | 15.1% | 2.63 | 2.27 | 1.15 | 49.1 |
| 5120 × 17408 | FFN down, 4-bit | 11.5% | 6.85 | 6.47 | 1.06 | 53.3 |
| 5120 × 17408 | FFN down, 5-bit | 9.2% | 7.36 | 6.46 | 1.14 | 49.6 |
| 10240 × 5120 | QKV projection, 4-bit | 8.8% | 3.92 | 3.59 | 1.09 | 54.7 |
| 248320 × 5120 | output vocabulary | 4.5% | 90.7 | 88.3 | 1.03 | 57.4 |
M=2048, bf16 activations, group size 64, transposed weights, chained dependent ops, median of 6 reps after warmup, idle GPU. MLX 0.32.3 on M5 Max (40-core GPU).
Every shape lands between 1.03× and 1.15× fp16. The 5-bit layers pay ~5% more than 4-bit for the messier dequant. The small-M pathology from #3584 does not appear at M=2048; it is a small-batch and speculative-decode problem, not a prefill one.
And the fp16 bars themselves tell the rest of the story. Plain matmul tops out at 57–62 TFLOPS on this chip. The one independent datapoint we have for the NAX ceiling is the author of MLX issue #3925 measuring 15.4 TFLOPS fp16 on a 10-core base M5; scaled to 40 cores that is ~62. So MLX's fused GEMM is at the hardware peak, and the quantized GEMM is within 15% of it. There is no tile-tuning win to find in the dense path.
Three things, all small, ranked by measured size:
| Opportunity | Evidence | Ceiling | Status |
|---|---|---|---|
| Attention efficiency at long context | D=256 fused NAX attention drops from 60 TFLOPS (2K keys) to 43 TFLOPS (16K keys) in isolation; 27% of the MoE step at 16K | ≤7% end to end | Youngest kernel here (PR #3842, weeks old). Worth a controlled tile experiment. |
| Split-K routing at tiny N | 1.5–1.7% of step on both models | ~1.5% | Already filed upstream (#3584, #4198). A dispatch-threshold PR, with correctness tests. |
| MoE expert GEMM at short runs | Not measured here; #3925 shows BM=64 tiles waste work when few tokens route to each expert | 6.6% TTFT on Qwen3.6-35B per the issue author | Upstream has a branch. Our gemma step is 27% expert GEMM; worth confirming on this box. |
None of these is "4×". The 4× already shipped; it is in the hardware and MLX turned it on. What this rules out matters as much as what it finds: if you are choosing where to spend engineering time to make an M5 feel faster for agents, do not spend it on MLX's dense kernels. Spend it above them, in the serving layer: prefix-cache restore paths, batching policy, and speculative decoding that actually composes with hybrid models. That is the next thing we will measure.
quantized.cpp and scaled_dot_product_attention.cpp falsified half the premise before any hardware was touched..gputrace is not a data format; it is an Xcode session. Op-level instrumentation with source-derived dispatch prediction got us 90% of the answer, headless, in an afternoon.Harness and plan: ~/hermes/m5-nax-coverage/ (nax_harness.py, p2_shapes.py, PLAN.md); will be pushed to GitHub with the engine-layer follow-up. Prior art credited: MLX issues #3584, #3925, #4198, PR #3842; Apple's MLX on M5 post; SiliconBench for the "serving software, not silicon" framing.