Three packs
All three keep the model's MTP head and run the same exact speculative path. They differ in quantization, which sets download size, memory, speed, and how close the output stays to the bf16 model. M1 and M2 Macs get an FP16 sibling of each pack automatically (same weights, native precision for chips without bf16).
| Pack | Quant | Download | Peak memory | KL vs bf16 | Pick it for |
|---|---|---|---|---|---|
| Optimized Speed | 4-bit dynamic, 32-weight groups, 8-bit on sensitive modules | 20.4 GB | 23.6 GB | 0.0220 | Coding. The default on 32 GB+ Macs. |
| Bare Speed | Flat 4-bit, 64-weight groups | 16.0 GB | 17.0 GB | 0.0376 | Fastest chat. Lower quality and slower on long coding tasks. |
| Optimized Quality | 8-bit dynamic, 64-weight groups | 29.4 GB | 32.7 GB | 0.00105 | Closest to the bf16 model. 36 GB+ Macs. |
KL divergence is measured against the original bf16 model on MTPLX's coding battery. Optimized Speed sits 1.7x closer to the original than Bare Speed; Optimized Quality sits 21x closer than Optimized Speed. A second measurement on 16 September 2026, teacher-forced over 2,389 positions of code, prose, JSON and a multilingual notice against the bf16 checkpoint: Optimized Speed agrees with bf16 on 96.0 percent of top-1 tokens with a KL of 0.012, Optimized Quality on 99.3 percent with a KL of 0.0005. For scale, mlx-serve's own score for its 4-bit 27B pack is 83.6 percent and KL 0.322, and 95.5 percent and 0.0136 for its 8-bit (different corpus, their method). MTPLX moved its flagship off a flat 4-bit trunk in July 2026 (Qwen 3.6 Optimized Speed V2) because the calibrated map held up better as agent sessions got longer, and flat 4-bit is slower on long coding tasks. The dynamic map below is why the recommended pack costs four more gigabytes than the bare one.
Measured speeds
M5 Max, fans verified at max, single stream, generation running to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20), MTPLX 2.7.0, 15 August 2026. Same prompt across the three packs.
| Run | Optimized Speed | Bare Speed | Optimized Quality |
|---|---|---|---|
| Coding task, medium reasoning, mtplx serve | 58.7 tok/s | 65.2 tok/s | 40.6 tok/s (depth 2) |
| Same task inside the Mac app, cold session | 55.5 tok/s | 64.4 tok/s | 48.3 tok/s |
| Long reasoning at xhigh, 20k to 46k-token answers | 35.1 to 37.3 tok/s | 32.0 to 35.7 tok/s | 33.1 to 33.2 tok/s |
| One 52,740-token answer, 27.2 minutes, model's own stop | 32.4 tok/s sustained | ||
| Draft acceptance by depth on the coding task | 0.95 / 0.88 / 0.80 | 0.95 / 0.86 / 0.78 | 0.96 / 0.88 / 0.79 |
Same night, same task, other engines: the previous MTPLX flagship Qwen 3.6 27B Optimized Speed V2 ran 59.9 to 60.1 tok/s; oMLX 0.5.7 serving its own Qwen 3.8 4-bit MTP quant ran 63.3 tok/s; LM Studio on the 52,740-token answer ran 17.40 tok/s against 32.4 here. mlx-serve publishes 73 tok/s for its own Qwen 3.8 27B 4-bit pack at temperature 0 on an M4 Max with MTP forced on (its benchmarks.md at v26.9.3), a different machine and sampler; see MTPLX vs mlx-serve.
Speed by context length matters more than any single number once an agent is in the loop. On MTPLX 2.10.0, stock settings, the Optimized Speed pack decodes a 3k-token chat answer at 64.3 tok/s, holds 30.4 tok/s at 88k context and 18.4 tok/s at 147k. The full set, with the release each was measured on, is on the benchmarks page.
MTPLX 2.11 (4 September 2026) took the 88k decode round apart with the kernel clock: of a 128 ms round, 38 ms
was attention in the verify kernel and 54 ms was weight streaming, which does not move with context. Attention
was the term that grew, so it got a flash-decoding verify kernel, sdpa_nax_flash, on by default in
the Turbo profile for this pack (MTPLX_NAX_FLASH_ROUTE=0 restores the old route). On the 72.7k walk
bench it runs 0.917 ms per layer against the previous kernel's 1.421 at half the power. Die-matched pairs on the
same prompts against the 2.10.2 engine, M5 Max, Optimized Speed, from the
2.11 release note:
| Context | Decode gain, 2.11 against the 2.10.2 engine | Notes |
|---|---|---|
| 16k tokens | +7% | the flash-decoding route, the shipped default |
| 88k tokens | +16% | the route alone, the shipped default |
| 88k tokens, two opt-in settings | +51% | MTPLX_CONTEXT_COPY=0 and batched target rows (MTPLX_LAZY_TARGET_DISTRIBUTIONS=0 MTPLX_BATCH_TARGET_ARRAYS=1); both are null or slightly negative at 16k, so they stay operator settings |
These pairs ran a harder prompt set than the 2.10.0 curve above, so the gains are the comparable figure; the pairs themselves are in the release note. The kernel changes only the order of the reduction; the acceptance rule is unchanged at every temperature. Draft depth 3 stays the optimum on this head; a depth of 4 keeps 3.1 tokens per round and adds 11 percent to the round.
RAM and Macs
- 32 GB or more: Optimized Speed. On M3, M4 and M5 Macs it is the app's first recommendation from 32 to 255 GB; from 256 GB the first recommendation is Flash-Next Optimized Speed.
- 36 GB or more: Optimized Quality fits with headroom.
- 16 to 31 GB: run Ternary Bonsai 2 27B instead. Since 2.12.0 it is the first recommendation on M3, M4 and M5 Macs with 16 to 31 GB, and MiMo V2.6 Qwen 9B the second; Qwen 3.5 9B and 4B fit too. Picked by hand on a 16 or 24 GB Mac, the Qwen 3.8 27B does not fit and gets the 4,096-token minimum window, with one startup line that says so. The app checks your Mac before recommending anything.
- M1 and M2: the app and CLI pick the FP16 build of the same pack automatically.
- Context window: 262,144 tokens nominal. The memory governor resolves the largest window whose KV fits and prints it in the serve banner. On 2.12.0 the 27B plans 57,344 tokens on 36 GB, 204,800 on 48 GB and 262,144 from 64 GB. On 2.10.0 a 48 GB Mac got 196,608 tokens, and at 42k context that seat measured 33 tok/s decode and 645 tok/s prefill. Since 2.11,
mtplx serve --allow-swap(or the app's Settings > Memory card) admits prompts past the memory fit for operators who accept swap, and an explicitMTPLX_MEMORY_LIMIT_BYTESis honored as the engine budget both ways.
Install
Mac app: download the DMG, pick "Qwen 3.8 27B Optimized Speed". It is the default on modern Macs. The app downloads the pack, sets up its engine, and measures your machine to pick the fastest decoding depth.
Command line:
brew install youssofal/mtplx/mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed
Then point OpenCode, Pi, Claude Code, Cline, Cursor or anything that speaks the OpenAI or Anthropic API at
http://127.0.0.1:8000. Setup pages for each client are in the docs. The tuned
depth and draft settings ship inside mtplx_runtime.json; MTPLX reads them on load, no flags needed.
Reasoning effort levels (xhigh, medium, low) work; MTPLX's coding default is medium.
How it is built
The Optimized Speed pack is a dynamic 4-bit quant with a hand-tuned precision map, the same layout as the Qwen 3.6 Optimized Speed V2 that preceded it:
- The bulk of the model at 4-bit with 32-weight groups.
- The parts that hurt most at 4-bit kept at 8-bit: the embeddings, the output head, all 48 GDN output projections, and the last 8 MLP blocks.
- The GDN convolution kernels and recurrent-state parameters, every norm, and the whole MTP head at 16-bit.
- Activations in bf16 on every pass, plain decode and speculative verify alike. The speculative path runs the same model as the autoregressive path.
Bare Speed is every weight matrix at 4-bit with 64-weight groups, nothing promoted. Optimized Quality is every weight matrix at 8-bit with 64-weight groups. All three keep the same 16-bit set.
Exactness
Speculation in MTPLX is exact. Drafts from the MTP head are accepted with the probability-ratio rule and rejected drafts are resampled from the residual, so what you sample is what the model would have sampled without speculation, at any temperature. The draft sampler is a speed knob only. The quantization is the one approximation, and the KL numbers above measure it.