MTPLX/Docs/Run Qwen 3.8 on a Mac

The fastest way to run Qwen 3.8 on a Mac.

MTPLX. On an M5 Max, Qwen 3.8 27B decodes up to 87.6 tok/s rewriting a file it just wrote (Optimized Speed, MTPLX 2.10.0, stock settings), 65.2 tok/s on a fresh medium-reasoning coding task (Bare Speed, MTPLX 2.7.0, official Qwen 3.8 sampling) and 64.3 tok/s on a 3k-token chat answer; Qwen 3.8 Flash-Next, the 125B MoE, runs up to 84 tok/s on warm agent turns (MTPLX 2.11). Qwen 3.8 ships with its own multi-token-prediction head and MTPLX uses it: the model drafts three tokens ahead, the target verifies them in one pass, and every accepted token is an exact sample from the model's own distribution. MTPLX supported Qwen 3.8 on 15 Aug 2026, the day after the model was released.

Pick the pack that fits your memory, install MTPLX, run mtplx start. The speeds to expect are listed with the conditions they were measured under, and the exactness check at the end confirms on your own Mac that the speedup leaves the output unchanged. The headline numbers on an M5 Max: Qwen 3.8 Flash Next at 125.8 tok/s on an OpenCode request and 79.3 tok/s at 9k tokens of context (MTPLX 2.11.3), Qwen 3.8 27B up to 87.6 tok/s. For scale: published Ollama runs of Qwen 3.8 27B Q4_K_M on other people's Macs report 10 to 17 tok/s (M1 Max to M4 Max, August 2026), against MTPLX's 58.7 to 65.2 tok/s on the same model class on an M5 Max. Created by Youssof Altoukhi, who brought native MTP to the Mac in April 2026.

Pick a pack by RAM

Qwen 3.8 27B comes in three MTPLX packs, all built from Qwen/Qwen3.8-27B with a 262,144-token context and MTP depth 3. The 125B Flash-Next MoE is a separate model for 96 GB and larger Macs. Each of the three 27B packs has an FP16 sibling for M1 and M2, the same weights in the precision those chips run natively; the app and CLI pick it automatically, and no M1 or M2 numbers are published.

PackQuantizationDownloadPeak memoryKL vs bf16RAM
Optimized Speed (recommended for coding)4-bit dynamic, 32-weight groups; 8-bit on embeddings, output head, all 48 GDN output projections and the last 8 MLP blocks; 16-bit GDN conv and recurrent-state params, norms and the whole MTP head20.4 GB23.6 GB0.022032 GB+ (default)
Bare Speedflat 4-bit, 64-weight groups16.0 GB17.0 GB0.0376no separate tier published
Optimized Quality8-bit, 64-weight groups29.4 GB32.7 GB0.0010536 GB+
Flash-Next Optimized Speed (125B MoE)dynamic 4-bit with 8-bit attention115.1 GB incl. the 32 GB n-gram table~83 GB resident weights plus working set96 GB+
Flash-Next Bare Speed (125B MoE)flat 4-bit96 GB+

Optimized Speed is the default on Macs with 32 GB or more and the pack to pick for coding. Bare Speed is described on its card as "Quickest burst chat speeds. Lower quality and slower on long coding tasks." Optimized Quality is the 8-bit pack for Macs with 36 GB or more. Below 32 GB, the Qwen 3.5 9B and 4B packs run comfortably on 16 GB. Quantization maps and every measured number are on the Qwen 3.8 27B page and the Flash-Next page.

Install and run

The Mac app from the DMG checks your hardware, recommends the pack that fits, downloads it, and measures your machine to pick the fastest decoding depth. From the terminal:

brew install youssofal/mtplx/mtplx
mtplx start

Or with pip, python3 -m pip install -U mtplx. To fetch a specific pack ahead of time and read its compatibility report before anything runs:

mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed
mtplx inspect Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed --json

Onboarding runs the model itself at each draft depth with fans pinned, keeps plain autoregressive decoding as the baseline, and saves a depth only if it beats that baseline. The Qwen 3.8 packs run the Turbo profile by default (2.0.0, 6 Jul 2026): verify-specialized quantized-matmul kernels plus a compiled verify step. The server then answers on 127.0.0.1:8000 for any OpenAI-compatible or Anthropic-compatible client; the quickstart covers the first request.

Expected speeds

All numbers below are from the public Hugging Face model cards, measured on an M5 Max with fans verified at max, single stream, generation running to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20), MTPLX 2.7.0, 15 Aug 2026.

PackCoding task, medium reasoningLong xhigh reasoning answersAcceptance by depthVerify stepConditions
Qwen 3.8 27B Optimized Speed58.7 tok/s via mtplx serve; 55.5 in the app35.1 tok/s on a 28k-token answer; 37.3 on a 20k-token answer0.95 / 0.88 / 0.8050 to 53 msMTPLX 2.7.0, 15 Aug 2026, M5 Max, fans verified at max, single stream, official Qwen 3.8 sampling
Qwen 3.8 27B Bare Speed65.2 tok/s via mtplx serve; 64.4 in the app35.7 tok/s on a 34k-token answer; 32.0 on a 37k-token answer; one 52,740-token answer in 27.2 minutes at 32.4 tok/s sustained0.95 / 0.86 / 0.7844 msMTPLX 2.7.0, 15 Aug 2026, M5 Max, fans verified at max, single stream, official Qwen 3.8 sampling
Qwen 3.8 27B Optimized Quality48.3 tok/s in the app33.2 tok/s on a 34k-token answer; 33.1 on a 46k-token answer0.96 / 0.88 / 0.7963.5 msMTPLX 2.7.0, 15 Aug 2026, M5 Max, fans verified at max, single stream, official Qwen 3.8 sampling

Two later instruments. The MTPLX 2.9.0 pack table for Qwen 3.8 27B (20 Aug 2026, depth 3, M5 Max, official Qwen 3.8 sampling; a different instrument from the 2.7.0 coding task) lists Optimized Speed at 46.8 tok/s, 2.3x the same model's plain decode without MTP; Bare Speed 49.9 tok/s, 2.3x; Optimized Quality 39.2 tok/s, 3.0x; Speed FP16 45.4, 2.3x; Bare FP16 50.2, 2.3x; Quality FP16 48.7, 2.8x. The 2.10.0 context curve (29 Aug 2026, M5 Max, stock settings, Qwen 3.8 27B Optimized Speed) is the one to read for long sessions:

Lane: Qwen 3.8 27B Optimized Speed, M5 Max, stock settings, MTPLX 2.9.2 vs 2.10.0 (29 Aug 2026)2.9.22.10.0
3k-token chat answer, decode55.9 tok/s64.3 tok/s
88k context, decode23.6 tok/s30.4 tok/s
147k context, decode12.0 tok/s18.4 tok/s
88k context, prefill379 tok/s535 tok/s
Rewriting a file the model just wrote (cache-copy rewrite lane)73.8 tok/s87.6 tok/s

MTPLX 2.11 (4 Sep 2026) adds a flash-decoding verify kernel for the 27B, on by default in the Turbo profile. Of a 128 ms decode round at 88k context, 38 ms was attention and 54 ms weight streaming; attention was the term that grew with context, so it got the new kernel. Die-matched pairs on the same prompts against the 2.10.2 engine, M5 Max, Qwen 3.8 27B Optimized Speed, on a harder prompt set than the 2.10.0 curve above: +7% decode at 16k context and +16% at 88k from the route alone; +51% at 88k with two opt-in long-context settings (MTPLX_CONTEXT_COPY=0 and batched target rows). Those two stay operator settings because both are null or slightly negative at 16k. The pairs themselves are in the 2.11 release note.

On the same night (15 Aug 2026, MTPLX 2.7.0, M5 Max, official Qwen 3.8 sampling, fans verified at max) the same coding task at medium reasoning ran on other engines and models: Qwen 3.6 27B Optimized Speed V2 on MTPLX 59.9 to 60.1 tok/s; oMLX 0.5.7 serving its own Qwen 3.8 4-bit MTP quant 63.3 tok/s, against MTPLX Qwen 3.8 27B Bare Speed at 65.2 and Optimized Speed at 58.7. On the 52,740-token long xhigh answer that night (same machine, same sampling), LM Studio ran 17.40 tok/s against MTPLX Qwen 3.8 27B Bare Speed at 32.4. oMLX is a continuous-batching server with SSD caching, and its Lightning MTP runs on MTPLX kernels. LM Studio's speculative decoding uses a separate draft model. The full set is on the benchmarks page.

Memory. The memory governor (2.10.0) sizes the context window to what your Mac can hold and prints the plan in the serve banner: a 48 GB Mac serving the 27B Speed pack resolves 196,608 tokens instead of the nominal 262,144, and at 42k context that seat measured 33 tok/s decode and 645 tok/s prefill. A request that cannot fit is refused up front with HTTP 507 (2.10.2, 1 Sep 2026) instead of dying mid-stream.

Flash Next on 96 GB+ Macs: 125 tok/s

MTPLX 2.11.3 (17 September 2026) on a MacBook Pro M5 Max with 128 GB, the Optimized Speed pack, fans verified at maximum, sampled at the model's own settings: 125.8 tok/s on one OpenCode request (1,301 tokens generated, 18,539-token prompt with 18,364 tokens served from cache), 54.6 to 122.4 tok/s per model request on real OpenCode coding tasks, 79.3 tok/s on a 9k-token code prompt (62.5 on 2.11.2, same Mac and runtime), 61.8 tok/s on a 109k-token OpenCode turn and 50.3 tok/s at 200k. A 96,760-token conversation restored from the session cache in 8 ms. Every row with its conditions is on the Flash Next page and the benchmarks page; the comparison with mlx-serve, using their own M5 Max table, is at MTPLX vs mlx-serve.

Qwen 3.8 Flash-Next is the 125B MoE. MTPLX 2.10.0 (29 Aug 2026) shipped the first Apple Silicon backend for the family, in two packs: Optimized Speed (dynamic 4-bit with 8-bit attention) and Bare Speed (flat 4-bit, the quickest). MTPLX 2.11 (4 Sep 2026) decodes it on an M5 Max with 128 GB at 68.4 tok/s at 16k context with a 1,024-token answer, 60.9 tok/s at 100k and 44.2 tok/s at 206k, temperature 1, copy lane on, coding prompts, same-hour pairs against 2.10.2 (53.2, 47.5 and 32.2). The lane behind that is a compiled verify path plus decode and prefill items adapted from PR #391 by @davidtai, ported under his name, with MTPLX's own fences and gates; every item is token-identical to the tree before it at temperature 0. The Optimized Speed download is 115.1 GB including a 32 GB n-gram table that streams from SSD; resident weights are about 83 GB plus the working set, which is why 96 GB is the floor. Clients may request depths up to 5.

2.10.1 (30 Aug 2026) added block-sparse prefill and image input. On the M5 Max a 98k-token Flash-Next prompt went from 175.7 s to 114.5 s with peak memory from 91.4 to 83.0 GB, and a cold 262,144-token prompt completes in 355 s at 87.4 GB peak, where it previously reached 119 GB or did not complete. 96 GB Macs can load Flash-Next; since 2.11 an explicit MTPLX_MEMORY_LIMIT_BYTES is honored as the engine budget, so a 96 GB Mac serving the Bare Speed pack under a limit of 80G runs instead of being refused. Warm agent turns on Flash-Next run 0.15 s to first token and 66 to 84 tok/s at 40k tokens of context on the 2.11 agent-session gate, and the first token after a pause of a few seconds takes 0.08 s of server time instead of 1.06 s.

Check exactness

MTPLX's claim is that the speculative path produces the same distribution of output as plain decoding of the same model. Acceptance is probability-ratio min(1, p/q) with residual (p - q)+ resampling (Leviathan–Chen); activations stay bf16 (fp16 on M1 and M2) on every pass, and the speculative path runs the same weights as the plain path. MTPLX has never shipped a greedy-only path. You can check this on your own Mac with the same loaded model.

Run the same prompt twice, once with generation_mode set to "ar" (plain decoding, the response reports mtp_depth: 0) and once with the default. Fix seed and set temperature to 0 for the first pass:

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"Write a Python function that parses ISO 8601 dates."}],"temperature":0,"seed":7,"generation_mode":"ar"}'

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"Write a Python function that parses ISO 8601 dates."}],"temperature":0,"seed":7}'

At temperature 0 the model's distribution has a single outcome, so the two replies should match token for token. Then repeat the pair at the family sampler (temperature 1.0, top-p 0.95, top-k 20). Now each reply is a draw from the same distribution and the wording may differ between the two; the guarantee is that over repeated runs neither path produces answers the other could not. The pack cards' KL numbers (0.0220 Optimized Speed, 0.0376 Bare Speed, 0.00105 Optimized Quality, each against bf16) measure the quantization; the speculative path samples from that same quantized model, so it adds nothing on top.

The same comparison is available at the server level and in the terminal chat: mtplx start --no-mtp serves plain decoding of the same loaded model, and inside mtplx start cli the commands /mtp off, /mtp on and /mtp status switch without reloading. /health reports mtp_enabled and depth so you know which path answered.

Since the day after release

Qwen released Qwen 3.8 on 14 Aug 2026. MTPLX 2.7.0 shipped Qwen 3.8 support on 15 Aug 2026, the day after, with the three 27B packs measured the same night. All six Qwen 3.8 27B repos ship their vision towers (restored 15 Aug 2026), so image input runs with MTP intact. On 3 Sep 2026 the Optimized Speed pack had 61,275 Hugging Face downloads in the trailing 30 days and all Qwen 3.8 packs together 119,875. The catalog lives at huggingface.co/Youssofal.