Conditions
All first-party numbers: single stream on a MacBook Pro M5 Max with 128 GB, fans pinned, sampled at the model's own sampler. Each row states the MTPLX version, pack, task or context length and date, and links to its release note or raw log. Measured by Youssof Altoukhi, who brought native MTP to the Mac.
Lanes
Decode speed on a laptop falls with context length, because every token reads the whole KV cache. A number is only comparable to another number in the same lane.
| Lane | What it is | Who it represents |
|---|---|---|
| Burst | A 192-token generation on a short coding prompt, thinking off (mtplx bench tune) | A short chat reply. Where speculative decoding looks best. |
| Chat | A 3k-token answer to a chat prompt | A normal conversation turn with reasoning. |
| Coding task | A medium-reasoning coding prompt, generation to the model's own stop | One agent turn writing code. |
| Long answer | Uncapped 11k to 52k-token answers | A full game or document in one turn; decay within one response. |
| Agent context | Decode at 88k and 147k tokens of context | An agent that has read files for an hour. Where most agent sessions run. |
| Rewrite | Rewriting a file the model just wrote, drafts copied from context | Edit-heavy agent turns. New-text decode speed is in the chat and agent-context rows. |
Qwen 3.8 Flash Next, the 125B MoE
MTPLX 2.12.0 against 2.11.3 on an M5 Max with 128 GB, 22 September 2026: the same Python runtime, boots in the order 2.11.3, 2.12.0, 2.12.0, 2.11.3, one model loaded at a time, fans at maximum, thinking off, 512-token answers, the model's own sampling with seed 1731, and the GPU memory limit raised to 120 GiB (the default is about 96 GiB). From the 2.12.0 release notes.
| Lane | 2.11.3 | 2.12.0 | Conditions | Source |
|---|---|---|---|---|
| Time to the first token, 4,061-token prompt | 5.26 s | 2.88 s | 22 Sep 2026 | 2.12.0 |
| Prompt processing, 4,061-token prompt | 786 | 1,453 | +85%, tok/s of prompt processing | 2.12.0 |
| Decode after a 4,061-token prompt | 74.4 | 74.1 | the same | 2.12.0 |
| Time to the first token, 65,502-token prompt | 85.5 s | 60.1 s | 25 seconds sooner | 2.12.0 |
| Prompt processing, 65,502-token prompt | 768 | 1,094 | +42%, tok/s of prompt processing | 2.12.0 |
| Decode after a 65,502-token prompt | 56.1 | 63.5 | +13%; at this length 2.11.3 copied the whole key and value cache on every verify round, and 2.12.0 writes it in place | 2.12.0 |
The five prompt kernels in 2.12.0 were added after that run, so the release is faster than those rows show. Their own paired run, on the same M5 Max with alternating boots of 2.12.0 without and with the kernels:
| Lane | Without the kernels | With the kernels | Conditions | Source |
|---|---|---|---|---|
| Prompt processing, 16,376-token prompt | 1,459 | 1,695 | +16%; time to the first token 11.3 s to 9.7 s | 2.12.0 |
| Prompt processing, 65,529-token prompt | 1,415 | 1,548 | +9%; time to the first token 46.6 s to 42.6 s; peak memory 92.0 GB with and without; logits and 256 greedy tokens bit-identical | 2.12.0 |
On M5 chips Flash-Next reads prompts 4,096 tokens at a time and switches to sparse attention from 16K tokens. M1 to M4 keep 2,048-token chunks and the quantized multiply, and their speed was not measured.
MTPLX 2.11.3 (17 September 2026), the Optimized Speed pack, MacBook Pro M5 Max with 128 GB, fans verified at maximum, sampled at the model's own settings. Every row is from the 2.11.3 release notes and the measurements behind them.
| Run | tok/s | Conditions | Source |
|---|---|---|---|
| One OpenCode request | 125.8 | 1,301 tokens generated, 18,539-token prompt with 18,364 tokens served from cache, MTP depth 3, OpenCode Desktop, 16 September 2026; the fastest measured request | 2.11.3 |
| OpenCode coding tasks, per model request | 54.6 to 122.4 | real multi-file tasks on the installed daemon, seven of eight requests served from the session cache; Pi 53.6 to 72.1, Hermes 56.1 to 82.4 | 2.11.3 |
| 9k-token code prompt | 79.3 | 1,500 tokens generated, seeded sampler, thinking off, two alternating boots each; 62.5 on 2.11.2, same Mac and runtime (+27 percent) | 2.11.3 |
| 109k-token OpenCode turn | 61.8 | mean of two runs, 60.8 and 62.7; 48.8 on the same build before the launcher and depth-policy fix, memory unchanged | 2.11.3 |
| 200k-token OpenCode turn, warm | 50.3 | mean of two runs, 50.0 and 50.6; in the installed app the cold 200,073-token prompt decoded at 42.5 and the warm follow-up at 50.6 | 2.11.3 |
| Full 45k to 56k-token generations | 66.8 | whole turn, the Flappy Bird prompt at effort xhigh; 66.5 and 62.6 on 2.11.2 on alternating boots of the same runtime | 2.11.3 |
| Rewrite after a 4 ms restore | 73.2 | installed app: a 54,530-token conversation restored in 4 ms, then a 42,200-token rewrite; the third turn restored 96,760 tokens in 8 ms | 2.11.3 |
| Exactness | at temperature 1, top-p 0.95, top-k 20, a thousand four-token draws from the fast path match a thousand from the plain path within the plain path's own split-half noise at every joint length, on Flash Next and on the 27B Quality pack | 2.11.3 |
MTPLX 2.11 against 2.10.2 on the same M5 Max with 128 GB: same-hour pairs, temperature 1, copy lane on in both arms, coding prompts. The lane is a compiled verify path plus decode and prefill items adapted from PR #391 by @davidtai, ported commit by commit under his name, with MTPLX's own fences and gates around it; every item is token-identical to the tree before it at temperature 0. Packs are described on the model page.
| Lane | 2.10.2 | 2.11 | Conditions | Source |
|---|---|---|---|---|
| 16k context, 1,024-token answer | 53.2 | 68.4 | +29%; round 48.2 to 39.2 ms; cold-prefill time to first token 15.1 to 14.3 s; peak memory 89.4 GB, flat; 4 Sep 2026 | 2.11 |
| 100k context | 47.5 | 60.9 | +28%; round 54.5 to 43.4 ms; cold-prefill time to first token 117.0 to 113.2 s; peak 97.9 GB against 98.1 GB | 2.11 |
| 206k context | 32.2 | 44.2 | +37%; at this rung the compiled lane is memory-gated on 128 GB, so 44.2 is the shipped default there | 2.11 |
| Warm agent turns, 40k context | 66 to 84 | the release's agent-session gate replayed against the live daemon: a long real-code prompt, short same-session turns, an auto tool round and a forced-choice round; 0.15 s to first token, 0.01 s dead time | 2.11 | |
| Cold page cache | 56 | 68.8 | decode after model load with the n-gram table pre-read (--ngram-prewarm auto) reading the hottest rows of the 30 GiB sidecar into the page cache | 2.11 |
| First token after a pause | 1.06 s | 0.08 s | server time to first token on "hi" after 3 to 90 s of quiet; first text on screen 0.87 to 0.26 s; the engine keeps its 77 GiB working set GPU-resident while attentive | 2.11 |
The launch numbers, from the pack cards and the 2.10.x release notes:
| Lane | Pack | tok/s | Conditions | Source |
|---|---|---|---|---|
| Coding task | Optimized Speed / Bare Speed | 73.5 / 75.9 | MTP against 43.8 / 47.0 plain autoregressive on the same task, mtplx serve, M5 Max, official Qwen 3.8 sampling, fans verified, single stream, 29 Aug 2026 | 2.10.0 |
| Prefill at 131k | Flash-Next | 810 | block-sparse prefill, tok/s of prefill; a 262,144-token cold prompt completes in 355 s at 87.4 GB peak, 30 Aug 2026 | 2.10.1 |
Qwen 3.8 27B, the current flagship for coding
Official Qwen 3.8 sampling throughout: temperature 1.0, top-p 0.95, top-k 20. Packs are described on the model page.
| Lane | Pack | tok/s | Conditions | Source |
|---|---|---|---|---|
| Coding task | Bare Speed | 65.2 | M5 Max, mtplx serve, medium reasoning, fans verified, to the model's stop, 15 Aug 2026 | 2.7.0 |
| Coding task | Optimized Speed | 58.7 | same run; acceptance by depth 0.95 / 0.88 / 0.80 | 2.7.0 |
| Coding task, in the app | Bare Speed | 64.4 | installed app, cold session, 17.0 GB peak | 2.7.0 |
| Coding task, in the app | Optimized Speed | 55.5 | installed app, cold session, 23.6 GB peak | 2.7.0 |
| Coding task, in the app | Optimized Quality (8-bit) | 48.3 | installed app, cold session, 32.7 GB peak | 2.7.0 |
| Long answer | Bare Speed | 32.4 | one 52,740-token answer, 27.2 minutes, ended at the model's own stop | 2.7.0 |
| Long answer | Optimized Speed | 35.1 to 37.3 | xhigh reasoning, 28k and 20k-token answers | 2.7.0 |
| Depth-3 pack table | Optimized Speed / Bare / Quality | 46.8 / 49.9 / 39.2 | 2.3x / 2.3x / 3.0x the same pack decoded without MTP, quantized draft heads, 20 Aug 2026 | 2.9.0 |
| Chat | Optimized Speed | 64.3 | 3k-token chat answer, stock settings, 29 Aug 2026 (55.9 on 2.9.2) | 2.10.0 |
| Agent context | Optimized Speed | 30.4 | 88k tokens of context (23.6 on 2.9.2) | 2.10.0 |
| Agent context | Optimized Speed | 18.4 | 147k tokens of context (12.0 on 2.9.2) | 2.10.0 |
| Rewrite | Optimized Speed | 87.6 | rewriting a file the model just wrote, cache-copy drafting (73.8 on 2.9.2) | 2.10.0 |
| Prefill | Optimized Speed | 535 | prefill tok/s at 88k context (379 on 2.9.2) | 2.10.0 |
| 48 GB seat | Optimized Speed | 33 | simulated 48 GB Mac, 42k context, 645 tok/s prefill, memory governor | 2.10.0 |
| Agent first token | Optimized Speed | 0.11 s | mid-session turn with preserved reasoning (1.8 to 2.2 s before) | 2.10.0 |
| Claude Code, 165k first turn | Optimized Speed | 3.9 s | follow-up time to first token, 165,165 of 165,502 tokens from cache, 1 Sep 2026 | 2.10.2 |
| Flash-decoding verify, 16k | Optimized Speed | +7% | decode gain from the new sdpa_nax_flash route, on by default in Turbo; die-matched pair on the same prompts against the 2.10.2 engine, a harder prompt set than the 2.10.0 curve, so the gain is the comparable figure and the pair is in the release note; 4 Sep 2026 | 2.11 |
| Flash-decoding verify, 88k | Optimized Speed | +16% | same pairing at 88k tokens of context from the route alone; +51% with two opt-in long-context settings, MTPLX_CONTEXT_COPY=0 and batched target rows, which are null or slightly negative at 16k | 2.11 |
Ternary Bonsai 2 27B against the 4-bit Qwen 3.8 27B
MTPLX 2.12.0, the same M5 Max with 128 GB, 22 September 2026: one model loaded at a time, temperature 1.0, top-p 0.95, top-k 20, thinking off, 512-token answers, and the 27B in the same session. Bonsai 2 is Prism ML's Qwen3.8-27B rebuilt with ternary weights, an 8.85 GB pack. Its "before" figures are alternating boots of 2.12.0 without its two new GPU kernels. The pack is described on the model page.
| Lane | Bonsai 2 27B | Qwen 3.8 27B Optimized Speed | Conditions | Source |
|---|---|---|---|---|
| Decode after a 4,061-token prompt | 64.4 | 52.6 | peak memory 11.4 GB against 23.9 GB; Bonsai 50.4 before the new kernels (+28%) | 2.12.0 |
| Decode after a 16,350-token prompt | 57.1 | 51.0 | peak memory 14.6 GB against 27.0 GB; Bonsai 45.2 before the new kernels (+26%) | 2.12.0 |
On the same Mac limited to the 16 GB class's 12.0 GiB engine budget, with the draft head active, a 7,006-token prompt and a 1,024-token answer peaked at 11.54 GiB of GPU memory and 12.79 GiB for the whole process, with no swap growth. This was measured on a 128 GB Mac under a 16 GB budget, not on a real 16 GB Mac.
Agent sessions
A 43k-token OpenCode session on Flash-Next, measured on the live daemon before MTPLX 2.11, spent 146 seconds of a 14-minute task waiting on the engine while its decode rounds ran 65 to 70 tok/s: 5 to 8 seconds before the first token of every tool turn, hidden waits on the previous turn's cache snapshot, and two whole-turn re-prefills of 38 and 18 seconds. 2.11 fixes the five engine defects behind that, and the release script now replays the loop against a serving daemon and fails on any of them. M5 Max, 128 GB, 4 September 2026, from the 2.11 release note.
| Measure | Before 2.11 | 2.11 | Conditions |
|---|---|---|---|
| Tool turn after a file write, tokens re-prefilled | 3,535 (3.7 s) | 20 (0.12 s) | the follow-up after a write in the 43k-token OpenCode session; the tool-call turn's final state is banked in place |
| Dead time before the first token of a warm tool turn | 5 to 7.5 s | 0.01 s | warm turns whose prefill is 20 to 400 tokens; the request no longer starts and then waits on an n-gram gather it will never use |
| Forced tool_choice round in a 41k-token session | 85 s | 0.58 s | two cold re-prefills of 41 and 44 s before; a per-request instruction no longer changes the session's cache identity |
| Agent-session gate at 40k tokens | 0.15 s / 66 to 84 tok/s | warm turns, time to first token and decode; the gate fails on dead time over 1 s, a warm first token over 1.5 s, a restore under 90 percent of the prompt, or a decode drop past 20 percent | |
| Block-restorable turn near the memory line | 54 s cold re-prefill | restores | a 41,901-token turn with a 41,391-token restore available; the pre-prefill memory guard now pins such entries instead of evicting them |
Decode by context length
Qwen 3.8 27B Optimized Speed, M5 Max, stock settings, MTPLX 2.9.2 against 2.10.0, from the 2.10.0 release note. Decode at 3k runs 3.5 times faster than at 147k on the same Mac; every engine has a curve like it.
Qwen 3.8 Flash-Next, M5 Max 128 GB, temperature 1, copy lane on, coding prompts, MTPLX 2.10.2 against 2.11, from the 2.11 release note.
Between turns the session cache matters more than peak decode speed: a 100k-token session restores in about two seconds after a restart instead of a five-minute cold prefill, and mid-session tool rounds restore warm in under two seconds (2.0.0).
Records
Records by run, each with the long-form number from the same configuration where one exists.
| Date | Record | tok/s | Conditions | Source |
|---|---|---|---|---|
| 16 Sep 2026 | 125B MoE, one OpenCode request | 125.8 | Qwen 3.8 Flash Next Optimized Speed, 1,301 tokens generated, 18,539-token prompt with 18,364 tokens from cache, MTP depth 3, M5 Max 128 GB, MTPLX 2.11.3 | 2.11.3 |
| 16 Sep 2026 | 125B MoE at 9k context | 79.3 | Qwen 3.8 Flash Next, 9k-token code prompt, 1,500 tokens generated, thinking off, two alternating boots each; 62.5 on 2.11.2 | 2.11.3 |
| 16 Sep 2026 | 125B MoE at 109k and 200k context | 61.8 / 50.3 | Qwen 3.8 Flash Next, 109k-token and 200k-token OpenCode turns, means of two runs each, M5 Max 128 GB | 2.11.3 |
| 2 Jul 2026 | The 27B record on a fresh generation | 81.74 | Qwen 3.6 27B Optimized Speed, depth 3, 192-token bench, thinking off, temp 0.6, fans verified 7,821 to 7,830 RPM, twin runs 81.74 / 81.73; plain decode 30.37 (2.69x) | raw logs |
| 2 Jul 2026 | Same config, long answer | 62.95 | uncapped 11,390-token Flappy Bird, reasoning on, seed 0, clean stop; 75.5 over the first 128 tokens, 39.9 over the last 128 | notes |
| 7 Jul 2026 | Qwen 3.5 9B, 6-bit verify kernels | 82.9 to 112.5 | Qwen 3.5 9B 6-bit, short context, M5 Max; 61.6 to 99.7 at 8k | 2.0.1 |
| 18 Jul 2026 | Qwen 3.5 4B, rebuilt draft head | 227.8 | Qwen 3.5 4B Optimized Speed, depth 3, 1.71x over 133.6 plain decode, M5 Max | 2.2.0 |
| 18 Jul 2026 | Qwen 3.6 35B-A3B, depth 2 default | 145 | Qwen 3.6 35B-A3B MoE at the new default depth 2 against 132 at the old depth 3, M5 Max | 2.2.0 |
| 29 Aug 2026 | Qwen 3.8 27B rewrite lane | 87.6 | Qwen 3.8 27B Optimized Speed rewriting a file it just wrote, cache-copy drafting, M5 Max, stock settings (73.8 on 2.9.2) | 2.10.0 |
| 4 Sep 2026 | 125B MoE, warm agent turns | 66 to 84 | Qwen 3.8 Flash-Next at 40k tokens of context on the 2.11 agent-session gate, live daemon, M5 Max 128 GB; 0.15 s to first token | 2.11 |
| 29 Aug 2026 | MTP on a 125B MoE | 63 to 76 | Qwen 3.8 Flash-Next through the MTPLX server, M5 Max, 61 plain decode, workload dependent | 2.10.0 |
| 30 Aug 2026 | Prefill at 131k | 810 | Flash-Next block-sparse prefill, 131k-token prompt, tok/s of prefill; 262,144 tokens cold in 355 s | 2.10.1 |
| 4 Sep 2026 | 125B MoE at 16k context | 68.4 | Qwen 3.8 Flash-Next, 16k-token coding prompt with a 1,024-token answer, M5 Max 128 GB, temperature 1, copy lane on, same-hour pair against 2.10.2 (53.2) | 2.11 |
| 4 Sep 2026 | 125B MoE at 206k context | 44.2 | Qwen 3.8 Flash-Next, 206k-token coding prompt, same conditions, against 2.10.2 (32.2) | 2.11 |
Memory
| Date | What | Peak | Conditions | Source |
|---|---|---|---|---|
| 6 May 2026 | Sustained Mode | 27.5 GB | Qwen 3.6 27B at 32k context on Ivan Fioravanti's M5 Max benchmark, down from 98.6 GB | History |
| 7 May 2026 | Long-context curve | 22.1 / 27.2 / 37.5 GB | 32k / 64k / 128k, Sustained, with 620.6 / 504.3 / 372.1 prefill tok/s and 39.1 / 31.2 / 25.3 decode | v0.2.0 |
| 6 Jul 2026 | Packed verify attention | −8 GB / −16 GB | at 64k and 128k context, Qwen 3.6 27B; decode at 128k from 17 to 20+ tok/s | 2.0.0 |
| 15 Aug 2026 | Qwen 3.8 27B packs | 17.0 / 23.6 / 32.7 GB | Bare Speed / Optimized Speed / Optimized Quality, coding task in the app | 2.7.0 |
| 29 Aug 2026 | Memory governor | 196,608 tokens | context window resolved for a 48 GB Mac on the 27B Speed pack instead of the nominal 262,144; allocator growth 8.6 GB to 0.6 GB | 2.10.0 |
| 30 Aug 2026 | 262k cold prompt | 87.4 GB | Flash-Next, block-sparse prefill, 355 s; 98k prompt peak 91.4 to 83.0 GB | 2.10.1 |
| 4 Sep 2026 | Flash-Next decode at 16k and 100k | 89.4 / 97.9 GB | peak on the 2.11 pairs, flat against 2.10.2 at 16k and 98.1 GB at 100k; the compiled verify lane adds about 28 KB of bank per context token, so a per-request gate hands roughly the last 20 percent of the window on 128 GB to the plain verify | 2.11 |
| 4 Sep 2026 | n-gram pre-read budget | 14.2 GiB | automatic page-cache pre-read of the Flash-Next n-gram table at model load on a 128 GB Mac, down from 23.4 GiB, the amount the page cache can still hold once a long prefill has grown the engine to its envelope | 2.11 |
Other engines
Same-machine runs against Ollama, LM Studio, llama.cpp, mlx-lm and oMLX are on the compare pages.
Notes
- One chip. Everything first-party is an M5 Max with 128 GB. The app measures your own Mac during onboarding and picks the depth that wins there; MTPLX has published no M1, M2, M3 or M4 numbers of its own. Community-measured numbers with the command line and the log are welcome in issues.
- Sampled, never greedy. Every number is sampled at the model's shipped sampler.