I went looking for M5 kernel fruit. Apple already ate it.

Created Last updated
Measurement note · James Meadlock & Milo (James's AI agent, this session on Claude Fable 5.1 by Anthropic, reasoning effort medium) · hardware: MacBook Pro M5 Max, 128 GB · MLX 0.32.3
The verdict: on an M5 Max, MLX already routes ~98% of LLM prefill compute through the Neural Accelerator kernels, and its 4-bit matmuls run within 3–15% of full fp16 speed at ~55–57 TFLOPS. There is no big, unowned kernel win hiding in the dense path. The remaining software headroom is small (long-context attention, a few percent) and lives above the kernels, in the serving engines. This is a negative result, and a useful one.
98%
of prefill matmul + attention time on NAX kernels
1.03–1.15×
4-bit matmul cost vs fp16, same shape
57 TFLOPS
fp16 GEMM on M5 Max, measured
845 tok/s
Qwen3.8-27B cold 16K prefill

Contents

  1. Why we looked
  2. What the Neural Accelerators are, in one diagram
  3. How we measured (and what did not work)
  4. Where 19 seconds of prefill go
  5. The test that settled it: 4-bit vs fp16
  6. What is actually left
  7. Lessons

Why we looked

The M5 generation added a matrix unit to every GPU core. Apple calls them Neural Accelerators; the MLX source calls the kernel family NAX. Apple's own numbers say prompt processing (prefill) got roughly 4× faster than M4. Prefill is exactly the part of local inference that hurts for agents, because an agent turn is mostly reading context: our production traffic runs about 58 prompt tokens for every completion token.

So the hypothesis was tempting: new silicon, young software, a compute-bound workload. Surely there are ops MLX is not routing through the new hardware yet, and surely fixing that is low-hanging fruit.

Before touching a kernel we did two things: read the MLX source, and ask a second model (GPT-6 Astra) to tear the plan apart. The source reading showed MLX already has NAX paths for dense GEMM, quantized GEMM, mixture-of-experts gather-GEMM, attention prefill, and gated delta-net. Astra's review said the plan was optimising the wrong number: a higher share of time on NAX kernels is neither necessary nor sufficient for lower latency. Measure latency, use coverage as an explanation. That reframing is the reason this post has a clean conclusion instead of a spreadsheet of kernel names.

What the Neural Accelerators are, in one diagram

An LLM forward pass is mostly matrix multiplies. During prefill, thousands of tokens hit each weight matrix at once, so the multiply is big and square-ish: compute-bound. During decode, one token at a time hits each matrix: a skinny multiply that is bound by memory bandwidth, not math. The accelerators help the first case and do nothing for the second. That is why Apple quotes 4× for prefill and ~1.2× for generation.

Prefill is compute-bound and uses the M5 Neural Accelerators; decode is bandwidth-bound and does not Two panels. Left: prefill, a large matrix times a large matrix, routed to NAX kernels. Right: decode, a single row times a large matrix, routed to bandwidth-bound vector kernels. Where the M5 Neural Accelerators help, and where they cannot PREFILL · reading the prompt thousands of tokens hit every weight matrix at once tokens 2048 x 5120 × weights 5120 x 17408 big × big = compute-bound NAX kernels: qmm_t_nax, steel_attention_nax DECODE · writing the answer one token at a time, every weight matrix streamed from memory 1 token 1 x 5120 × weights 5120 x 17408 skinny × big = bandwidth-bound vector kernels: qmv, sdpa_vector (no NAX) Apple: ~4x faster prefill, ~1.2x faster decode on M5 vs M4. The shapes explain the asymmetry.

How we measured (and what did not work)

The obvious tool is Metal frame capture: mx.metal.start_capture() writes a .gputrace that Xcode's GPU Frame Debugger can replay with per-kernel timings. We got that far. A 482-token prefill on a 4B model produced a 3.8 GB bundle of 1,053 files. But the bundle is not scriptable: dispatches reference pipeline objects by hash, kernel names live in a 190 MB embedded metallib of everything MLX compiled (not what it ran), and there are no timings at all until Xcode replays the capture. Capturing also slows the run 10× (100 ms to 1,060 ms), so captured runs are never timing data. Astra had flagged exactly this risk. It was right.

So we went one level up. The harness monkey-patches the four MLX ops that carry the compute (quantized_matmul, gather_qmm, matmul, fast.scaled_dot_product_attention, plus gated_delta_update for hybrid models), records the shapes of every call, times each call with a GPU synchronize, and predicts which kernel MLX dispatched by applying the gating rules copied from the MLX source (quantized.cpp, matmul.cpp, scaled_dot_product_attention.cpp, main at 7916d8b). Those predictions are labelled predicted until someone opens one capture in Xcode and checks them; we have not done that yet.

How MLX decides which quantized matmul kernel to run Decision tree from the MLX source: small M goes to vector kernels; transposed non-batched shapes may take a split-K path with no NAX variant; aligned shapes on M5 take the NAX kernel. How MLX picks a quantized-matmul kernel (from quantized.cpp) M = tokens in this step · N, K = weight matrix sides · the harness applies these rules to predict each dispatch quantized_matmul(x, W) shape M × K times K × N Is M below the vector limit? (13–33 on M5) yes: decode, small batch qmv / qmv_wide bandwidth kernel, no NAX by design no: prefill Transposed, one batch, and split-K heuristic > 1? yes: tiny N qmm_t_splitk no NAX variant (issues #3584, #4198) we saw it only at N=32/48: 1.5% of step M5 GPU and K % 64 == 0 and not fp32? yes (every real model shape we hit) qmm_t_nax (bm32 if M ≤ 32, else bm64) Neural Accelerator tensor kernel 92% of Qwen3.8-27B prefill time lands here no: unaligned / fp32 qmm (legacy tiles) never observed on these models Attention has its own gate: head dim in {64, 96, 128, 256, 512} reaches steel_attention_nax; 72/80 are zero-padded to 96.

Where 19 seconds of prefill go

Two models already on the machine, chosen for the dense/MoE contrast: Qwen3.8-27B (dense, a mixed 4/5-bit MLX quant) and gemma-4-26b-a4b (mixture of experts, 4-bit). Workload: a 16,384-token synthetic agent prompt, cold, prefilled in mlx_lm's default 2,048-token chunks; then a 512-token delta appended to the warm cache, which is what an agent turn looks like after a tool call. Wall times below are uninstrumented; the breakdown comes from a second, instrumented pass.

ModelCold 16K prefilltok/sWarm 512 deltaTime on NAX kernels (predicted)
Qwen3.8-27B dense19.4 s845828 ms98.5%
gemma-4-26b-a4b MoE4.3 s3,850416 ms98.3%
Share of prefill compute time by kernel, both models Stacked horizontal bars. Qwen3.8-27B: 92 percent quantized matmul on NAX, 6.6 percent attention on NAX, 1.5 percent split-K without NAX. gemma-4-26b MoE: 44.6 percent quantized matmul NAX, 27.1 percent attention NAX, 26.6 percent expert gather matmul NAX, 1.7 percent split-K. Cold 16K prefill: where the compute time goes share of synchronized per-op time, matmul + attention family only · green = Neural Accelerator kernel · amber = no NAX variant Qwen3.8-27B dense · 19.4 s quantized matmul (qmm_t_nax) 92.0% attention 6.6% split-K 1.5% gemma-4-26b MoE · 4.3 s qmm_t_nax 44.6% attention 27.1% experts 26.6% 1.7% dense quantized GEMM, NAX attention prefill, NAX MoE expert GEMM, NAX split-K, no NAX kernel Takeaway: the only non-accelerated slice is 1.5–1.7% of the step. Coverage is not the problem.

The amber slice is real and already known upstream: for very narrow outputs (N=32 or 48, the per-head gate projections in Qwen3.5-style models) MLX's split-K heuristic picks a kernel that has no NAX variant. Issues #3584 and #4198 cover it with microbenchmarks. On a real model at real prefill sizes it is 1.5% of the step. Fixing it is worth a PR; it is not worth a project.

The test that settled it: 4-bit vs fp16

If 92% of the time is in one kernel family, the question becomes whether that kernel is good, not whether it is used. The cleanest test: take the exact matrix shapes the 27B model hits during a 2,048-token step, run each one through the quantized kernel and through plain fp16 matmul, and compare. A quantized kernel has to dequantize on the fly, so it should be a little slower than fp16 per FLOP. If it were much slower, that would be the fruit. Issue #3584 reports exactly that pathology at small M (1.5–1.8×).

Shape (N × K)What it isShare of 27B matmul time4-bitfp16ratio4-bit TFLOPS
17408 × 5120FFN up/gate, 4-bit37.5%6.48 ms5.92 ms1.0956.3
5120 × 6144attention out, 5-bit15.1%2.632.271.1549.1
5120 × 17408FFN down, 4-bit11.5%6.856.471.0653.3
5120 × 17408FFN down, 5-bit9.2%7.366.461.1449.6
10240 × 5120QKV projection, 4-bit8.8%3.923.591.0954.7
248320 × 5120output vocabulary4.5%90.788.31.0357.4

M=2048, bf16 activations, group size 64, transposed weights, chained dependent ops, median of 6 reps after warmup, idle GPU. MLX 0.32.3 on M5 Max (40-core GPU).

Quantized matmul throughput versus fp16 on the model's real shapes Paired horizontal bars for six matrix shapes. fp16 reaches 56 to 62 TFLOPS; 4-bit quantized reaches 49 to 57 TFLOPS. Ratios range from 1.03 to 1.15. The quantized kernel is already at the fp16 ceiling TFLOPS on the six shapes that carry 87% of Qwen3.8-27B prefill matmul time · longer is better · scale 0–65 0 30 TFLOPS 60 17408 × 5120 FFN up 4-bit 1.09× 5120 × 6144 attn out 5-bit 1.15× 5120 × 17408 FFN down 4-bit 1.06× 5120 × 17408 FFN down 5-bit 1.14× 10240 × 5120 QKV 4-bit 1.09× 248320 × 5120 vocab 4-bit 1.03× fp16 matmul (steel_gemm_fused_nax) 4/5-bit quantized_matmul (qmm_t_nax)

Every shape lands between 1.03× and 1.15× fp16. The 5-bit layers pay ~5% more than 4-bit for the messier dequant. The small-M pathology from #3584 does not appear at M=2048; it is a small-batch and speculative-decode problem, not a prefill one.

And the fp16 bars themselves tell the rest of the story. Plain matmul tops out at 57–62 TFLOPS on this chip. The one independent datapoint we have for the NAX ceiling is the author of MLX issue #3925 measuring 15.4 TFLOPS fp16 on a 10-core base M5; scaled to 40 cores that is ~62. So MLX's fused GEMM is at the hardware peak, and the quantized GEMM is within 15% of it. There is no tile-tuning win to find in the dense path.

Back of the envelope. A ~27B-parameter dense model does roughly 2 × 27B × 16,384 ≈ 885 TFLOP of matmul work on a 16K prompt. At 55 TFLOPS that is ~16 s. We measured 19.4 s end to end. The remaining ~3 s is attention, norms, rotary embeddings, elementwise ops, and dispatch gaps. Prefill on an M5 Max is compute-bound at the silicon ceiling, and the software is delivering about 80% of it.

What is actually left

Three things, all small, ranked by measured size:

OpportunityEvidenceCeilingStatus
Attention efficiency at long contextD=256 fused NAX attention drops from 60 TFLOPS (2K keys) to 43 TFLOPS (16K keys) in isolation; 27% of the MoE step at 16K≤7% end to endYoungest kernel here (PR #3842, weeks old). Worth a controlled tile experiment.
Split-K routing at tiny N1.5–1.7% of step on both models~1.5%Already filed upstream (#3584, #4198). A dispatch-threshold PR, with correctness tests.
MoE expert GEMM at short runsNot measured here; #3925 shows BM=64 tiles waste work when few tokens route to each expert6.6% TTFT on Qwen3.6-35B per the issue authorUpstream has a branch. Our gemma step is 27% expert GEMM; worth confirming on this box.

None of these is "4×". The 4× already shipped; it is in the hardware and MLX turned it on. What this rules out matters as much as what it finds: if you are choosing where to spend engineering time to make an M5 feel faster for agents, do not spend it on MLX's dense kernels. Spend it above them, in the serving layer: prefix-cache restore paths, batching policy, and speculative decoding that actually composes with hybrid models. That is the next thing we will measure.

What we did not do. No Xcode replay of a capture yet, so kernel attribution is predicted from source rules, not observed. No engine-level comparison (mlx_lm vs oMLX vs vllm-metal) yet. Two models, one chip. BaseRT (arXiv 2607.19438) reports up to 3.9× prefill over MLX on an M5 Pro with hand-written kernels; we cite it as related work, not as a ceiling for this box, because its engine binary is closed and its workloads differ. Per-op timing with a synchronize after each call is a ranking tool, not an absolute attribution: the instrumented sum ran within ±20% of the uninstrumented wall time on both models.

Lessons

Reproduce

Harness and plan: ~/hermes/m5-nax-coverage/ (nax_harness.py, p2_shapes.py, PLAN.md); will be pushed to GitHub with the engine-layer follow-up. Prior art credited: MLX issues #3584, #3925, #4198, PR #3842; Apple's MLX on M5 post; SiliconBench for the "serving software, not silicon" framing.