MTPLX/Models

MTP-ready models
for Apple Silicon.

Nine models, every pack shipped with the draft head MTPLX needs for exact speculative decoding on Apple Silicon. The Qwen packs keep the model's own native multi-token-prediction head; the Bonsai 2 pack adds the Qwen3.8-27B draft head, the MiMo pack uses the Qwen3.5-9B one because Xiaomi's checkpoint has none, and the Gemma 4 pack pairs the target with Google's official assistant drafter. Pick by RAM. The app checks your Mac before recommending anything, and auto-tune measures each draft depth on your machine during onboarding, keeps plain AR as the baseline, and saves a depth only if it beats AR. Created by Youssof Altoukhi, who brought native MTP to the Mac in April 2026.

Nine model pages
32 GB+ · Default for coding

Qwen 3.8 27B

Three packs. Up to 87.6 tok/s rewriting a file it just wrote (Optimized Speed, M5 Max, stock settings, MTPLX 2.10.0), 65.2 tok/s on a fresh medium-reasoning coding task (Bare Speed, mtplx serve, official Qwen 3.8 sampling, MTPLX 2.7.0) and 64.3 tok/s on a 3k-token chat answer. MTPLX 2.11 adds a flash-decoding verify kernel, +7% at 16k and +16% at 88k.

96 GB+ · 125B MoE

Qwen 3.8 Flash-Next

The fastest way to run Qwen 3.8 Flash Next on a Mac. MTPLX 2.11.3 on an M5 Max with 128 GB: 125.8 tok/s on an OpenCode request, 79.3 tok/s at 9k tokens of context (62.5 on 2.11.2), 61.8 tok/s on a 109k-token OpenCode turn, 50.3 at 200k, sampled at the model's own settings. First Apple Silicon backend for the family (MTPLX 2.10.0). 96 GB Macs and up, and since 2.12.0 an 8-bit Optimized Quality pack for Macs with 256 GB or more.

16 GB+ · 27B-class model

Ternary Bonsai 2 27B

Prism ML's Qwen3.8-27B rebuilt with ternary weights, with image input: an 8.85 GB pack that runs on 16 GB Macs. On an M5 Max with MTPLX 2.12.0 it decoded at 64.4 tok/s after a 4,061-token prompt against 52.6 for the 4-bit Qwen 3.8 27B, with a peak of 11.4 GB against 23.9 (temperature 1.0, thinking off, 512-token answers).

32 GB+ · Previous flagship

Qwen 3.6 27B

The model MTPLX was built on and the holder of the 27B record: 81.74 tok/s at 2.69x plain decode (Optimized Speed, depth 3, 192-token coding bench, thinking off, temperature 0.6, fans verified, M5 Max, 2 July 2026, raw logs published). On the medium-reasoning coding task shared with the Qwen 3.8 packs (MTPLX 2.7.0), Optimized Speed V2 ran 59.9 to 60.1 tok/s.

MoE · Sustained profile

Qwen 3.6 35B-A3B

145 tok/s at the default depth 2 against 132 at depth 3 (M5 Max, MTPLX 2.2.0). 4-bit MLX body with calibrated INT4 MTP heads. Since MTPLX 2.6.0 two agents at once decode at 1.6 to 2.25x per lane against the previous AR batch route on an M5 Max, sampled at shipped settings.

16 GB+ · Coding and agents

MiMo V2.6 Qwen 9B

Xiaomi's coding and agent fine-tune of Qwen 3.5 9B, as an 8.70 GB 6-bit pack with the Qwen 3.5 9B draft head. Xiaomi's model card reports 44.6 on SWE Pro against 32.0 for Qwen 3.5 9B. New in MTPLX 2.12.0, and the second suggestion on M3, M4 and M5 Macs with 16 to 31 GB, after Bonsai 2.

16 GB · Small-Mac pick

Qwen 3.5 9B

6-bit body, bf16 MTP head. MTPLX 2.0.1 (7 July 2026) took decode from 82.9 to 112.5 tok/s at short context and from 61.6 to 99.7 at 8k on an M5 Max with its 6-bit verify kernels.

8 GB+ · Fastest pack

Qwen 3.5 4B

2.47 GB download. 227.8 tok/s at depth 3 against 133.6 tok/s plain AR (1.71x) on an M5 Max, fans at max, MTPLX 2.2.0, the card's deterministic suite, 18 July 2026.

Assistant-pair build

Gemma 4 31B

Gemma 4 31B IT at MLX 4-bit paired with Google's official Gemma 4 31B assistant drafter at MLX 6-bit, verified with the same exact speculative sampling. In the catalog since MTPLX 1.0.0 (11 June 2026).

Pick by RAM

What fits your Mac.

Only the tiers MTPLX has published. Where no tier is published the app decides from the machine it runs on, and the memory governor (2.10.0) prints engine budget, weights, resolved context window and session bank in the serve banner. Requests that cannot fit are refused up front with HTTP 507 (2.10.2). Since 2.11 the app's Settings > Memory card sets a memory limit in GB and an allow-swap switch, and mtplx serve --allow-swap does the same for the CLI.

Unified memoryRunPublished numbers
8 GB or moreQwen 3.5 4B2.47 GB download, about 2.9 GiB peak at load. Runs on any Apple Silicon Mac with 8 GB or more.
16 to 31 GBTernary Bonsai 2 27B, then MiMo V2.6 Qwen 9B; Qwen 3.5 9B or 4BMTPLX 2.12.0 suggests Bonsai 2 first and MiMo second on M3, M4 and M5 Macs; M1 and M2 keep their FP16 recommendations. Bonsai 2: 8.85 GB download, an 8,192-token window on 16 GB, 20,480 on 18 GB and 94,208 on 24 GB, so agent clients such as OpenCode need 18 GB or more. MiMo: 8.70 GB download, 20,480 tokens on 16 GB. The README states that 16 GB of memory runs the 4B and 9B models comfortably.
32 GB or moreQwen 3.8 27B Optimized Speed, the first recommendation from 32 to 255 GB; Qwen 3.6 27B Optimized Speed V23.8 Optimized Speed: 20.4 GB download, 23.6 GB measured peak. 3.6 V2: 19.9 GB download.
36 GB or moreQwen 3.8 27B Optimized Quality8-bit. 29.4 GB download, 32.7 GB measured peak.
96 GB or moreQwen 3.8 Flash-Next: Bare Speed from 96 GB, Optimized Speed from 128 GBBare Speed peaks at 78 GiB and Optimized Speed at 87 GiB; resident weights about 74 GB and 83 GB plus working set, with the 32 GB n-gram table streaming from SSD. 125.8 tok/s on an OpenCode request, 79.3 at 9k context, 61.8 at 109k (MTPLX 2.11.3, M5 Max 128 GB).
256 GB or moreQwen 3.8 Flash-Next Optimized Speed first, then Optimized Quality (8-bit)Optimized Quality: 169.96 GB download, about 128.5 GiB of weights in memory with the n-gram table on SSD. It is listed second because it has not run on a 256 GB Mac yet, and its speed has not been measured (MTPLX 2.12.0).
No tier publishedQwen 3.6 35B-A3B, Gemma 4The app checks your Mac before recommending anything; auto-tune measures the speed on it.

M1 and M2 Macs get FP16 siblings of the Qwen 3.8 27B, Qwen 3.6 35B-A3B and Qwen 3.5 9B packs (same weights, native precision for chips without bf16). Every tok/s number on these pages carries its MTPLX version, pack, chip, task, sampler and date; the benchmarks page collects them. Speculation is exact on every pack: drafts are accepted with the probability-ratio rule and rejected drafts are resampled from the residual, so the output follows the model's distribution at any temperature. The acceptance math is on how it works.

Forge converts a Hugging Face repo into an MTP-ready MLX build on your own Mac and measures the speedup; it refuses incompatible models instead of silently falling back. Hugging Face 30-day downloads on 3 September 2026: 61,275 for Qwen 3.8 27B Optimized Speed, 119,875 across the Qwen 3.8 packs, 138,089 across all MTPLX packs. All packs live under huggingface.co/Youssofal.