MTPLX/Benchmarks

125 tok/s on Flash Next,
81.74 on a 27B, on a MacBook.

The MTPLX record lane: Qwen 3.6 27B Optimized Speed at 81.74 tok/s, 2.69x plain decode, on a MacBook Pro M5 Max, depth 3, thinking off, temperature 0.6, fans verified, twin runs, raw logs below. Qwen 3.8 27B runs up to 87.6 tok/s rewriting a file it just wrote and 65.2 tok/s on a fresh coding task at official Qwen 3.8 sampling. Qwen 3.8 Flash-Next, a 125B MoE, runs up to 84 tok/s on warm agent turns and 68.4 tok/s at 16k context on MTPLX 2.11. Qwen 3.5 4B decodes 227.8 tok/s and Qwen 3.5 9B 112.5 tok/s. On MTPLX 2.12.0, Flash-Next reads a 4,061-token prompt 85 percent faster than 2.11.3, and Ternary Bonsai 2 27B decodes faster than the 4-bit Qwen 3.8 27B in about half the memory. Every row below carries its version, pack, chip, context and date.

Conditions

All first-party numbers: single stream on a MacBook Pro M5 Max with 128 GB, fans pinned, sampled at the model's own sampler. Each row states the MTPLX version, pack, task or context length and date, and links to its release note or raw log. Measured by Youssof Altoukhi, who brought native MTP to the Mac.

Lanes

Decode speed on a laptop falls with context length, because every token reads the whole KV cache. A number is only comparable to another number in the same lane.

LaneWhat it isWho it represents
BurstA 192-token generation on a short coding prompt, thinking off (mtplx bench tune)A short chat reply. Where speculative decoding looks best.
ChatA 3k-token answer to a chat promptA normal conversation turn with reasoning.
Coding taskA medium-reasoning coding prompt, generation to the model's own stopOne agent turn writing code.
Long answerUncapped 11k to 52k-token answersA full game or document in one turn; decay within one response.
Agent contextDecode at 88k and 147k tokens of contextAn agent that has read files for an hour. Where most agent sessions run.
RewriteRewriting a file the model just wrote, drafts copied from contextEdit-heavy agent turns. New-text decode speed is in the chat and agent-context rows.

Qwen 3.8 Flash Next, the 125B MoE

MTPLX 2.12.0 against 2.11.3 on an M5 Max with 128 GB, 22 September 2026: the same Python runtime, boots in the order 2.11.3, 2.12.0, 2.12.0, 2.11.3, one model loaded at a time, fans at maximum, thinking off, 512-token answers, the model's own sampling with seed 1731, and the GPU memory limit raised to 120 GiB (the default is about 96 GiB). From the 2.12.0 release notes.

Lane2.11.32.12.0ConditionsSource
Time to the first token, 4,061-token prompt5.26 s2.88 s22 Sep 20262.12.0
Prompt processing, 4,061-token prompt7861,453+85%, tok/s of prompt processing2.12.0
Decode after a 4,061-token prompt74.474.1the same2.12.0
Time to the first token, 65,502-token prompt85.5 s60.1 s25 seconds sooner2.12.0
Prompt processing, 65,502-token prompt7681,094+42%, tok/s of prompt processing2.12.0
Decode after a 65,502-token prompt56.163.5+13%; at this length 2.11.3 copied the whole key and value cache on every verify round, and 2.12.0 writes it in place2.12.0

The five prompt kernels in 2.12.0 were added after that run, so the release is faster than those rows show. Their own paired run, on the same M5 Max with alternating boots of 2.12.0 without and with the kernels:

LaneWithout the kernelsWith the kernelsConditionsSource
Prompt processing, 16,376-token prompt1,4591,695+16%; time to the first token 11.3 s to 9.7 s2.12.0
Prompt processing, 65,529-token prompt1,4151,548+9%; time to the first token 46.6 s to 42.6 s; peak memory 92.0 GB with and without; logits and 256 greedy tokens bit-identical2.12.0

On M5 chips Flash-Next reads prompts 4,096 tokens at a time and switches to sparse attention from 16K tokens. M1 to M4 keep 2,048-token chunks and the quantized multiply, and their speed was not measured.

MTPLX 2.11.3 (17 September 2026), the Optimized Speed pack, MacBook Pro M5 Max with 128 GB, fans verified at maximum, sampled at the model's own settings. Every row is from the 2.11.3 release notes and the measurements behind them.

Runtok/sConditionsSource
One OpenCode request125.81,301 tokens generated, 18,539-token prompt with 18,364 tokens served from cache, MTP depth 3, OpenCode Desktop, 16 September 2026; the fastest measured request2.11.3
OpenCode coding tasks, per model request54.6 to 122.4real multi-file tasks on the installed daemon, seven of eight requests served from the session cache; Pi 53.6 to 72.1, Hermes 56.1 to 82.42.11.3
9k-token code prompt79.31,500 tokens generated, seeded sampler, thinking off, two alternating boots each; 62.5 on 2.11.2, same Mac and runtime (+27 percent)2.11.3
109k-token OpenCode turn61.8mean of two runs, 60.8 and 62.7; 48.8 on the same build before the launcher and depth-policy fix, memory unchanged2.11.3
200k-token OpenCode turn, warm50.3mean of two runs, 50.0 and 50.6; in the installed app the cold 200,073-token prompt decoded at 42.5 and the warm follow-up at 50.62.11.3
Full 45k to 56k-token generations66.8whole turn, the Flappy Bird prompt at effort xhigh; 66.5 and 62.6 on 2.11.2 on alternating boots of the same runtime2.11.3
Rewrite after a 4 ms restore73.2installed app: a 54,530-token conversation restored in 4 ms, then a 42,200-token rewrite; the third turn restored 96,760 tokens in 8 ms2.11.3
Exactnessat temperature 1, top-p 0.95, top-k 20, a thousand four-token draws from the fast path match a thousand from the plain path within the plain path's own split-half noise at every joint length, on Flash Next and on the 27B Quality pack2.11.3

MTPLX 2.11 against 2.10.2 on the same M5 Max with 128 GB: same-hour pairs, temperature 1, copy lane on in both arms, coding prompts. The lane is a compiled verify path plus decode and prefill items adapted from PR #391 by @davidtai, ported commit by commit under his name, with MTPLX's own fences and gates around it; every item is token-identical to the tree before it at temperature 0. Packs are described on the model page.

Lane2.10.22.11ConditionsSource
16k context, 1,024-token answer53.268.4+29%; round 48.2 to 39.2 ms; cold-prefill time to first token 15.1 to 14.3 s; peak memory 89.4 GB, flat; 4 Sep 20262.11
100k context47.560.9+28%; round 54.5 to 43.4 ms; cold-prefill time to first token 117.0 to 113.2 s; peak 97.9 GB against 98.1 GB2.11
206k context32.244.2+37%; at this rung the compiled lane is memory-gated on 128 GB, so 44.2 is the shipped default there2.11
Warm agent turns, 40k context66 to 84the release's agent-session gate replayed against the live daemon: a long real-code prompt, short same-session turns, an auto tool round and a forced-choice round; 0.15 s to first token, 0.01 s dead time2.11
Cold page cache5668.8decode after model load with the n-gram table pre-read (--ngram-prewarm auto) reading the hottest rows of the 30 GiB sidecar into the page cache2.11
First token after a pause1.06 s0.08 sserver time to first token on "hi" after 3 to 90 s of quiet; first text on screen 0.87 to 0.26 s; the engine keeps its 77 GiB working set GPU-resident while attentive2.11

The launch numbers, from the pack cards and the 2.10.x release notes:

LanePacktok/sConditionsSource
Coding taskOptimized Speed / Bare Speed73.5 / 75.9MTP against 43.8 / 47.0 plain autoregressive on the same task, mtplx serve, M5 Max, official Qwen 3.8 sampling, fans verified, single stream, 29 Aug 20262.10.0
Prefill at 131kFlash-Next810block-sparse prefill, tok/s of prefill; a 262,144-token cold prompt completes in 355 s at 87.4 GB peak, 30 Aug 20262.10.1

Qwen 3.8 27B, the current flagship for coding

Official Qwen 3.8 sampling throughout: temperature 1.0, top-p 0.95, top-k 20. Packs are described on the model page.

LanePacktok/sConditionsSource
Coding taskBare Speed65.2M5 Max, mtplx serve, medium reasoning, fans verified, to the model's stop, 15 Aug 20262.7.0
Coding taskOptimized Speed58.7same run; acceptance by depth 0.95 / 0.88 / 0.802.7.0
Coding task, in the appBare Speed64.4installed app, cold session, 17.0 GB peak2.7.0
Coding task, in the appOptimized Speed55.5installed app, cold session, 23.6 GB peak2.7.0
Coding task, in the appOptimized Quality (8-bit)48.3installed app, cold session, 32.7 GB peak2.7.0
Long answerBare Speed32.4one 52,740-token answer, 27.2 minutes, ended at the model's own stop2.7.0
Long answerOptimized Speed35.1 to 37.3xhigh reasoning, 28k and 20k-token answers2.7.0
Depth-3 pack tableOptimized Speed / Bare / Quality46.8 / 49.9 / 39.22.3x / 2.3x / 3.0x the same pack decoded without MTP, quantized draft heads, 20 Aug 20262.9.0
ChatOptimized Speed64.33k-token chat answer, stock settings, 29 Aug 2026 (55.9 on 2.9.2)2.10.0
Agent contextOptimized Speed30.488k tokens of context (23.6 on 2.9.2)2.10.0
Agent contextOptimized Speed18.4147k tokens of context (12.0 on 2.9.2)2.10.0
RewriteOptimized Speed87.6rewriting a file the model just wrote, cache-copy drafting (73.8 on 2.9.2)2.10.0
PrefillOptimized Speed535prefill tok/s at 88k context (379 on 2.9.2)2.10.0
48 GB seatOptimized Speed33simulated 48 GB Mac, 42k context, 645 tok/s prefill, memory governor2.10.0
Agent first tokenOptimized Speed0.11 smid-session turn with preserved reasoning (1.8 to 2.2 s before)2.10.0
Claude Code, 165k first turnOptimized Speed3.9 sfollow-up time to first token, 165,165 of 165,502 tokens from cache, 1 Sep 20262.10.2
Flash-decoding verify, 16kOptimized Speed+7%decode gain from the new sdpa_nax_flash route, on by default in Turbo; die-matched pair on the same prompts against the 2.10.2 engine, a harder prompt set than the 2.10.0 curve, so the gain is the comparable figure and the pair is in the release note; 4 Sep 20262.11
Flash-decoding verify, 88kOptimized Speed+16%same pairing at 88k tokens of context from the route alone; +51% with two opt-in long-context settings, MTPLX_CONTEXT_COPY=0 and batched target rows, which are null or slightly negative at 16k2.11
Other engines, same night, same task, same sampling (15 Aug 2026): Qwen 3.6 27B Optimized Speed V2 on MTPLX 59.9 to 60.1 tok/s; oMLX 0.5.7 serving its own Qwen 3.8 4-bit MTP quant 63.3 tok/s; LM Studio on the 52,740-token long answer 17.40 tok/s against 32.4 here.

Ternary Bonsai 2 27B against the 4-bit Qwen 3.8 27B

MTPLX 2.12.0, the same M5 Max with 128 GB, 22 September 2026: one model loaded at a time, temperature 1.0, top-p 0.95, top-k 20, thinking off, 512-token answers, and the 27B in the same session. Bonsai 2 is Prism ML's Qwen3.8-27B rebuilt with ternary weights, an 8.85 GB pack. Its "before" figures are alternating boots of 2.12.0 without its two new GPU kernels. The pack is described on the model page.

LaneBonsai 2 27BQwen 3.8 27B Optimized SpeedConditionsSource
Decode after a 4,061-token prompt64.452.6peak memory 11.4 GB against 23.9 GB; Bonsai 50.4 before the new kernels (+28%)2.12.0
Decode after a 16,350-token prompt57.151.0peak memory 14.6 GB against 27.0 GB; Bonsai 45.2 before the new kernels (+26%)2.12.0

On the same Mac limited to the 16 GB class's 12.0 GiB engine budget, with the draft head active, a 7,006-token prompt and a 1,024-token answer peaked at 11.54 GiB of GPU memory and 12.79 GiB for the whole process, with no swap growth. This was measured on a 128 GB Mac under a 16 GB budget, not on a real 16 GB Mac.

Agent sessions

A 43k-token OpenCode session on Flash-Next, measured on the live daemon before MTPLX 2.11, spent 146 seconds of a 14-minute task waiting on the engine while its decode rounds ran 65 to 70 tok/s: 5 to 8 seconds before the first token of every tool turn, hidden waits on the previous turn's cache snapshot, and two whole-turn re-prefills of 38 and 18 seconds. 2.11 fixes the five engine defects behind that, and the release script now replays the loop against a serving daemon and fails on any of them. M5 Max, 128 GB, 4 September 2026, from the 2.11 release note.

MeasureBefore 2.112.11Conditions
Tool turn after a file write, tokens re-prefilled3,535 (3.7 s)20 (0.12 s)the follow-up after a write in the 43k-token OpenCode session; the tool-call turn's final state is banked in place
Dead time before the first token of a warm tool turn5 to 7.5 s0.01 swarm turns whose prefill is 20 to 400 tokens; the request no longer starts and then waits on an n-gram gather it will never use
Forced tool_choice round in a 41k-token session85 s0.58 stwo cold re-prefills of 41 and 44 s before; a per-request instruction no longer changes the session's cache identity
Agent-session gate at 40k tokens0.15 s / 66 to 84 tok/swarm turns, time to first token and decode; the gate fails on dead time over 1 s, a warm first token over 1.5 s, a restore under 90 percent of the prompt, or a decode drop past 20 percent
Block-restorable turn near the memory line54 s cold re-prefillrestoresa 41,901-token turn with a 41,391-token restore available; the pre-prefill memory guard now pins such entries instead of evicting them

Decode by context length

Qwen 3.8 27B Optimized Speed, M5 Max, stock settings, MTPLX 2.9.2 against 2.10.0, from the 2.10.0 release note. Decode at 3k runs 3.5 times faster than at 147k on the same Mac; every engine has a curve like it.

3k · 2.10.064.3 tok/s
3k · 2.9.255.9 tok/s
88k · 2.10.030.4 tok/s
88k · 2.9.223.6 tok/s
147k · 2.10.018.4 tok/s
147k · 2.9.212.0 tok/s

Qwen 3.8 Flash-Next, M5 Max 128 GB, temperature 1, copy lane on, coding prompts, MTPLX 2.10.2 against 2.11, from the 2.11 release note.

16k · 2.1168.4 tok/s
16k · 2.10.253.2 tok/s
100k · 2.1160.9 tok/s
100k · 2.10.247.5 tok/s
206k · 2.1144.2 tok/s
206k · 2.10.232.2 tok/s

Between turns the session cache matters more than peak decode speed: a 100k-token session restores in about two seconds after a restart instead of a five-minute cold prefill, and mid-session tool rounds restore warm in under two seconds (2.0.0).

Records

Records by run, each with the long-form number from the same configuration where one exists.

DateRecordtok/sConditionsSource
16 Sep 2026125B MoE, one OpenCode request125.8Qwen 3.8 Flash Next Optimized Speed, 1,301 tokens generated, 18,539-token prompt with 18,364 tokens from cache, MTP depth 3, M5 Max 128 GB, MTPLX 2.11.32.11.3
16 Sep 2026125B MoE at 9k context79.3Qwen 3.8 Flash Next, 9k-token code prompt, 1,500 tokens generated, thinking off, two alternating boots each; 62.5 on 2.11.22.11.3
16 Sep 2026125B MoE at 109k and 200k context61.8 / 50.3Qwen 3.8 Flash Next, 109k-token and 200k-token OpenCode turns, means of two runs each, M5 Max 128 GB2.11.3
2 Jul 2026The 27B record on a fresh generation81.74Qwen 3.6 27B Optimized Speed, depth 3, 192-token bench, thinking off, temp 0.6, fans verified 7,821 to 7,830 RPM, twin runs 81.74 / 81.73; plain decode 30.37 (2.69x)raw logs
2 Jul 2026Same config, long answer62.95uncapped 11,390-token Flappy Bird, reasoning on, seed 0, clean stop; 75.5 over the first 128 tokens, 39.9 over the last 128notes
7 Jul 2026Qwen 3.5 9B, 6-bit verify kernels82.9 to 112.5Qwen 3.5 9B 6-bit, short context, M5 Max; 61.6 to 99.7 at 8k2.0.1
18 Jul 2026Qwen 3.5 4B, rebuilt draft head227.8Qwen 3.5 4B Optimized Speed, depth 3, 1.71x over 133.6 plain decode, M5 Max2.2.0
18 Jul 2026Qwen 3.6 35B-A3B, depth 2 default145Qwen 3.6 35B-A3B MoE at the new default depth 2 against 132 at the old depth 3, M5 Max2.2.0
29 Aug 2026Qwen 3.8 27B rewrite lane87.6Qwen 3.8 27B Optimized Speed rewriting a file it just wrote, cache-copy drafting, M5 Max, stock settings (73.8 on 2.9.2)2.10.0
4 Sep 2026125B MoE, warm agent turns66 to 84Qwen 3.8 Flash-Next at 40k tokens of context on the 2.11 agent-session gate, live daemon, M5 Max 128 GB; 0.15 s to first token2.11
29 Aug 2026MTP on a 125B MoE63 to 76Qwen 3.8 Flash-Next through the MTPLX server, M5 Max, 61 plain decode, workload dependent2.10.0
30 Aug 2026Prefill at 131k810Flash-Next block-sparse prefill, 131k-token prompt, tok/s of prefill; 262,144 tokens cold in 355 s2.10.1
4 Sep 2026125B MoE at 16k context68.4Qwen 3.8 Flash-Next, 16k-token coding prompt with a 1,024-token answer, M5 Max 128 GB, temperature 1, copy lane on, same-hour pair against 2.10.2 (53.2)2.11
4 Sep 2026125B MoE at 206k context44.2Qwen 3.8 Flash-Next, 206k-token coding prompt, same conditions, against 2.10.2 (32.2)2.11

Memory

DateWhatPeakConditionsSource
6 May 2026Sustained Mode27.5 GBQwen 3.6 27B at 32k context on Ivan Fioravanti's M5 Max benchmark, down from 98.6 GBHistory
7 May 2026Long-context curve22.1 / 27.2 / 37.5 GB32k / 64k / 128k, Sustained, with 620.6 / 504.3 / 372.1 prefill tok/s and 39.1 / 31.2 / 25.3 decodev0.2.0
6 Jul 2026Packed verify attention−8 GB / −16 GBat 64k and 128k context, Qwen 3.6 27B; decode at 128k from 17 to 20+ tok/s2.0.0
15 Aug 2026Qwen 3.8 27B packs17.0 / 23.6 / 32.7 GBBare Speed / Optimized Speed / Optimized Quality, coding task in the app2.7.0
29 Aug 2026Memory governor196,608 tokenscontext window resolved for a 48 GB Mac on the 27B Speed pack instead of the nominal 262,144; allocator growth 8.6 GB to 0.6 GB2.10.0
30 Aug 2026262k cold prompt87.4 GBFlash-Next, block-sparse prefill, 355 s; 98k prompt peak 91.4 to 83.0 GB2.10.1
4 Sep 2026Flash-Next decode at 16k and 100k89.4 / 97.9 GBpeak on the 2.11 pairs, flat against 2.10.2 at 16k and 98.1 GB at 100k; the compiled verify lane adds about 28 KB of bank per context token, so a per-request gate hands roughly the last 20 percent of the window on 128 GB to the plain verify2.11
4 Sep 2026n-gram pre-read budget14.2 GiBautomatic page-cache pre-read of the Flash-Next n-gram table at model load on a 128 GB Mac, down from 23.4 GiB, the amount the page cache can still hold once a long prefill has grown the engine to its envelope2.11

Other engines

Same-machine runs against Ollama, LM Studio, llama.cpp, mlx-lm and oMLX are on the compare pages.

Notes

  • One chip. Everything first-party is an M5 Max with 128 GB. The app measures your own Mac during onboarding and picks the depth that wins there; MTPLX has published no M1, M2, M3 or M4 numbers of its own. Community-measured numbers with the command line and the log are welcome in issues.
  • Sampled, never greedy. Every number is sampled at the model's shipped sampler.