MTPLX/Models/Qwen 3.8 27B

The fastest way to run Qwen 3.8 27B on a Mac.

The MTPLX flagship for coding on Macs with 32 GB or more. Qwen3.8-27B with its native multi-token-prediction head kept, so MTPLX drafts three tokens ahead and verifies them in one pass, with output identical in distribution to plain decoding. On an M5 Max it runs up to 87.6 tok/s rewriting a file it just wrote (Optimized Speed, MTPLX 2.10.0, stock settings), 65.2 tok/s on the launch-day coding task at Qwen's official sampling (Bare Speed, MTPLX 2.7.0, 15 August 2026, the day after the model was released; 58.7 on Optimized Speed) and 64.3 tok/s on a 3k-token chat answer (Optimized Speed, MTPLX 2.10.0). The record for a 27B on a Mac is 81.74 tok/s, set on Qwen 3.6 27B with the raw logs published. MTPLX 2.11 adds a flash-decoding verify kernel, +7% at 16k and +16% at 88k; 2.11.3 fixed eight exactness defects and measures the 4-bit pack at 96.0 percent top-1 agreement with bf16.

Three packs

All three keep the model's MTP head and run the same exact speculative path. They differ in quantization, which sets download size, memory, speed, and how close the output stays to the bf16 model. M1 and M2 Macs get an FP16 sibling of each pack automatically (same weights, native precision for chips without bf16).

On a 16 GB Mac, or to use less memory: Ternary Bonsai 2 27B, Prism ML's rebuild of Qwen3.8-27B with ternary weights, decodes faster than the 4-bit Optimized Speed pack in about half the memory. On MTPLX 2.12.0 it ran 64.4 against 52.6 tok/s after a 4,061-token prompt and 57.1 against 51.0 after a 16,350-token prompt, with a peak of 11.4 GB against 23.9 GB at the shorter prompt (M5 Max, temperature 1.0, top-p 0.95, top-k 20, thinking off, 512-token answers). It is an 8.85 GB download and runs on Macs with 16 GB.
PackQuantDownloadPeak memoryKL vs bf16Pick it for
Optimized Speed4-bit dynamic, 32-weight groups, 8-bit on sensitive modules20.4 GB23.6 GB0.0220Coding. The default on 32 GB+ Macs.
Bare SpeedFlat 4-bit, 64-weight groups16.0 GB17.0 GB0.0376Fastest chat. Lower quality and slower on long coding tasks.
Optimized Quality8-bit dynamic, 64-weight groups29.4 GB32.7 GB0.00105Closest to the bf16 model. 36 GB+ Macs.

KL divergence is measured against the original bf16 model on MTPLX's coding battery. Optimized Speed sits 1.7x closer to the original than Bare Speed; Optimized Quality sits 21x closer than Optimized Speed. A second measurement on 16 September 2026, teacher-forced over 2,389 positions of code, prose, JSON and a multilingual notice against the bf16 checkpoint: Optimized Speed agrees with bf16 on 96.0 percent of top-1 tokens with a KL of 0.012, Optimized Quality on 99.3 percent with a KL of 0.0005. For scale, mlx-serve's own score for its 4-bit 27B pack is 83.6 percent and KL 0.322, and 95.5 percent and 0.0136 for its 8-bit (different corpus, their method). MTPLX moved its flagship off a flat 4-bit trunk in July 2026 (Qwen 3.6 Optimized Speed V2) because the calibrated map held up better as agent sessions got longer, and flat 4-bit is slower on long coding tasks. The dynamic map below is why the recommended pack costs four more gigabytes than the bare one.

Measured speeds

M5 Max, fans verified at max, single stream, generation running to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20), MTPLX 2.7.0, 15 August 2026. Same prompt across the three packs.

RunOptimized SpeedBare SpeedOptimized Quality
Coding task, medium reasoning, mtplx serve58.7 tok/s65.2 tok/s40.6 tok/s (depth 2)
Same task inside the Mac app, cold session55.5 tok/s64.4 tok/s48.3 tok/s
Long reasoning at xhigh, 20k to 46k-token answers35.1 to 37.3 tok/s32.0 to 35.7 tok/s33.1 to 33.2 tok/s
One 52,740-token answer, 27.2 minutes, model's own stop32.4 tok/s sustained
Draft acceptance by depth on the coding task0.95 / 0.88 / 0.800.95 / 0.86 / 0.780.96 / 0.88 / 0.79

Same night, same task, other engines: the previous MTPLX flagship Qwen 3.6 27B Optimized Speed V2 ran 59.9 to 60.1 tok/s; oMLX 0.5.7 serving its own Qwen 3.8 4-bit MTP quant ran 63.3 tok/s; LM Studio on the 52,740-token answer ran 17.40 tok/s against 32.4 here. mlx-serve publishes 73 tok/s for its own Qwen 3.8 27B 4-bit pack at temperature 0 on an M4 Max with MTP forced on (its benchmarks.md at v26.9.3), a different machine and sampler; see MTPLX vs mlx-serve.

Speed by context length matters more than any single number once an agent is in the loop. On MTPLX 2.10.0, stock settings, the Optimized Speed pack decodes a 3k-token chat answer at 64.3 tok/s, holds 30.4 tok/s at 88k context and 18.4 tok/s at 147k. The full set, with the release each was measured on, is on the benchmarks page.

MTPLX 2.11 (4 September 2026) took the 88k decode round apart with the kernel clock: of a 128 ms round, 38 ms was attention in the verify kernel and 54 ms was weight streaming, which does not move with context. Attention was the term that grew, so it got a flash-decoding verify kernel, sdpa_nax_flash, on by default in the Turbo profile for this pack (MTPLX_NAX_FLASH_ROUTE=0 restores the old route). On the 72.7k walk bench it runs 0.917 ms per layer against the previous kernel's 1.421 at half the power. Die-matched pairs on the same prompts against the 2.10.2 engine, M5 Max, Optimized Speed, from the 2.11 release note:

ContextDecode gain, 2.11 against the 2.10.2 engineNotes
16k tokens+7%the flash-decoding route, the shipped default
88k tokens+16%the route alone, the shipped default
88k tokens, two opt-in settings+51%MTPLX_CONTEXT_COPY=0 and batched target rows (MTPLX_LAZY_TARGET_DISTRIBUTIONS=0 MTPLX_BATCH_TARGET_ARRAYS=1); both are null or slightly negative at 16k, so they stay operator settings

These pairs ran a harder prompt set than the 2.10.0 curve above, so the gains are the comparable figure; the pairs themselves are in the release note. The kernel changes only the order of the reduction; the acceptance rule is unchanged at every temperature. Draft depth 3 stays the optimum on this head; a depth of 4 keeps 3.1 tokens per round and adds 11 percent to the round.

RAM and Macs

  • 32 GB or more: Optimized Speed. On M3, M4 and M5 Macs it is the app's first recommendation from 32 to 255 GB; from 256 GB the first recommendation is Flash-Next Optimized Speed.
  • 36 GB or more: Optimized Quality fits with headroom.
  • 16 to 31 GB: run Ternary Bonsai 2 27B instead. Since 2.12.0 it is the first recommendation on M3, M4 and M5 Macs with 16 to 31 GB, and MiMo V2.6 Qwen 9B the second; Qwen 3.5 9B and 4B fit too. Picked by hand on a 16 or 24 GB Mac, the Qwen 3.8 27B does not fit and gets the 4,096-token minimum window, with one startup line that says so. The app checks your Mac before recommending anything.
  • M1 and M2: the app and CLI pick the FP16 build of the same pack automatically.
  • Context window: 262,144 tokens nominal. The memory governor resolves the largest window whose KV fits and prints it in the serve banner. On 2.12.0 the 27B plans 57,344 tokens on 36 GB, 204,800 on 48 GB and 262,144 from 64 GB. On 2.10.0 a 48 GB Mac got 196,608 tokens, and at 42k context that seat measured 33 tok/s decode and 645 tok/s prefill. Since 2.11, mtplx serve --allow-swap (or the app's Settings > Memory card) admits prompts past the memory fit for operators who accept swap, and an explicit MTPLX_MEMORY_LIMIT_BYTES is honored as the engine budget both ways.

Install

Mac app: download the DMG, pick "Qwen 3.8 27B Optimized Speed". It is the default on modern Macs. The app downloads the pack, sets up its engine, and measures your machine to pick the fastest decoding depth.

Command line:

brew install youssofal/mtplx/mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

Then point OpenCode, Pi, Claude Code, Cline, Cursor or anything that speaks the OpenAI or Anthropic API at http://127.0.0.1:8000. Setup pages for each client are in the docs. The tuned depth and draft settings ship inside mtplx_runtime.json; MTPLX reads them on load, no flags needed. Reasoning effort levels (xhigh, medium, low) work; MTPLX's coding default is medium.

How it is built

The Optimized Speed pack is a dynamic 4-bit quant with a hand-tuned precision map, the same layout as the Qwen 3.6 Optimized Speed V2 that preceded it:

  • The bulk of the model at 4-bit with 32-weight groups.
  • The parts that hurt most at 4-bit kept at 8-bit: the embeddings, the output head, all 48 GDN output projections, and the last 8 MLP blocks.
  • The GDN convolution kernels and recurrent-state parameters, every norm, and the whole MTP head at 16-bit.
  • Activations in bf16 on every pass, plain decode and speculative verify alike. The speculative path runs the same model as the autoregressive path.

Bare Speed is every weight matrix at 4-bit with 64-weight groups, nothing promoted. Optimized Quality is every weight matrix at 8-bit with 64-weight groups. All three keep the same 16-bit set.

Exactness

Speculation in MTPLX is exact. Drafts from the MTP head are accepted with the probability-ratio rule and rejected drafts are resampled from the residual, so what you sample is what the model would have sampled without speculation, at any temperature. The draft sampler is a speed knob only. The quantization is the one approximation, and the KL numbers above measure it.

Vision. All six Qwen 3.8 27B repos ship their vision towers (restored 15 August 2026 after the first upload went out text-only). Image input runs with MTP intact.