Start the server and launch Pi. The model runs on your Mac, the sampler is the model's own, and since 2.0.2 (9 Jul 2026) Pi's tool turns restore from the warm prefix cache. Created by Youssof Altoukhi, who brought native MTP to the Mac in April 2026.
One command
Install MTPLX from the DMG or with Homebrew, then:
brew install youssofal/mtplx/mtplx
mtplx start pi
mtplx start pi loads the model, serves it on 127.0.0.1:8000, and launches Pi against
the local server. The app's one-click Pi launch does the same thing. The app and the CLI share one server, so
mtplx start attaches to a model the app already has loaded instead of loading a second copy.
Manual setup
Pi uses the OpenAI-compatible endpoint, so a manual setup needs the same three values as any OpenAI-compatible client:
| Field | Value |
|---|---|
| Base URL | http://127.0.0.1:8000/v1 |
| Model | mtplx-qwen38-27b-optimized-speed |
| API key | any non-empty string for localhost |
The served id follows the loaded pack; /v1/models lists it. If you bind the server to a
non-localhost address you must start it with --api-key, and the server accepts that key as
Authorization: Bearer or X-API-Key.
curl http://127.0.0.1:8000/v1/models
The model's own sampling contract
MTPLX serves each pack at its own sampler. For Qwen 3.8 that is the official Qwen 3.8 sampling: temperature
1.0, top-p 0.95, top-k 20. Acceptance of drafted tokens is probability-ratio min(1, p/q) with
residual (p - q)+ resampling (Leviathan–Chen), so the output follows the model's distribution at
that sampler and at any other temperature. MTPLX has never shipped a greedy-only path. Activations stay bf16
(fp16 on M1 and M2) on every pass, and the speculative path runs the same model as the plain path.
A request may still set temperature, top_p, top_k,
presence_penalty and frequency_penalty; the server honors them per request. Penalties
default to 0, which is an exact no-op that preserves MTP exactness. Qwen's guidance is to leave them at 0 for
coding and agent work.
Warm tool turns
Since 2.0.0 (6 Jul 2026) MTPLX checkpoints the attention KV cache plus the recurrent and conv GDN state at
commit boundaries, with speculation on. 2.0.2 (9 Jul 2026) extended the warm prefix restore from OpenCode to Pi,
Claude Code and Cline. A Pi tool round that appends a result and resends the transcript restores the prefix
instead of prefilling it again: mid-session tool rounds restore warm in under 2 s, and a 100k-token session
restores in about 2 s after a restart instead of a five-minute cold prefill. Every response reports
usage.cached_tokens, so you can see what was served from cache.
MTPLX 2.11 (4 Sep 2026) removed the dead time between tool turns. Measured on a 43k-token agent session on
Flash-Next against the live daemon (M5 Max, 128 GB): the follow-up after a file write re-prefills 20 tokens in
0.12 s instead of 3,535 tokens in 3.7 s, the wait before a warm turn's first token is 0.01 s instead of 5 to
7.5 s, and a forced tool_choice round in a 41k-token session takes 0.58 s instead of 85 s. The fixes
are in the engine, so Pi's tool rounds ride the same path.
max_tokens at its model's contextWindow
minus a chars/4 estimate of the transcript and sends max_tokens=1 when the estimate overflows; the
answer then stops after one token, often inside a tool call. Since 2.11 the server logs one WARNING naming the
client's cap, and the generation trace line carries max_tokens, effective_max_tokens and
finish_reason. The fix on Pi's side is a larger contextWindow (#436).
MTPLX also reads Pi's models.json as JSONC since 2.11 and leaves a file it cannot parse untouched.Reasoning effort
In 2.10.0 (29 Aug 2026) OpenCode and Pi began defaulting to the family's agent-lane reasoning effort, medium
for Flash-Next, instead of xhigh. On the multi-file agent task in that release note, wall clock went from 150.2 s to 44.2 s with the same correct output; the gain came from the lower reasoning-effort default rather than from decode speed. The dial lives on the MTPLX side, in the app or through mtplx settings get and
mtplx settings set.
mtplx serve, M5 Max, fans
verified at max, single stream, official Qwen 3.8 sampling) Qwen 3.8 27B Optimized Speed decoded at 58.7 tok/s.
Flash-Next on MTPLX 2.11 (4 Sep 2026, M5 Max 128 GB, temperature 1, coding prompts): 68.4 tok/s at 16k context,
60.9 at 100k, 44.2 at 206k. The full set is on the benchmarks page.Community launcher
KrossKinetic/pi-mtplx on GitHub is a community
launcher that runs Pi against MTPLX. It is maintained outside the MTPLX repository; the built-in path is
mtplx start pi.
Model ids
The served id follows the pack.
| Pack | Served id | RAM |
|---|---|---|
| Qwen 3.8 27B Optimized Speed (recommended for coding) | mtplx-qwen38-27b-optimized-speed | 32 GB+ |
| Qwen 3.8 27B Optimized Quality (8-bit) | mtplx-qwen38-27b-optimized-quality | 36 GB+ |
| Qwen 3.8 Flash-Next Optimized Speed (125B MoE) | mtplx-flash-next-optimized-speed | 96 GB+ |
| Qwen 3.5 9B Optimized Speed | mtplx-qwen35-9b-optimized-speed | 16 GB+ |
Details for each pack, including the quantization map and the measured speeds, are on the models pages.