AX ENGINE / Performance
Performance.
AX Engine leads the loadable MTP peers on the Qwen 3.8 27B AXQ 6-bit MTP pack: MTPLX (draft depth 3), OMLX (Lightning depth 1), and the direct-AR `mlx-lm` baseline. Numbers are 20-run medians on the same checkpoint.
Mac mini M4 Pro 64 GB · 273 GB/s memory bandwidth
Qualification SKU · Mac mini M4 Pro 64 GB · 2026-09-17
| Runtime | Decode (tok/s) | Prefill (tok/s) |
|---|---|---|
| AX Engine 7.4.0 · ad999f3f | 31.05 | 120.3 |
| MTPLX 2.11.3 | 28.16 | 114.0 |
| oMLX 0.6.4 | 15.02 | — |
| mlx-lm 0.31.3 | 12.78 | — |
Apple M5 Max 128 GB · 614 GB/s memory bandwidth
Campaign host · Apple M5 Max 128 GB · 2026-09-15 (pre-correction)
| Runtime | Decode (tok/s) | Prefill (tok/s) |
|---|---|---|
| AX Engine 7.4.0 | 76.90 | 795.3 |
| MTPLX 2.11.2 | 70.62 | 686.6 |
| oMLX 0.6.4 | 38.47 | — |
| mlx-lm 0.31.3 | 27.90 | — |
MTP Tier 2 is still pending. Decode moves past the DRAM ceiling by emitting several tokens per weight pass; the bandwidth ceiling on the qualification SKU is reached at ~266 GB/s of weight-stream bandwidth.
Full peer comparison, methodology, and Tiel / Cyber-Tiel MXFP4 MTP rows ↗Qwen 3.8 27B benchmarks ↗