MTPLX/Models/Qwen 3.8 Flash-Next

The fastest way to run Qwen 3.8 Flash Next on a Mac.

Qwen 3.8 Flash Next (written Qwen3.8-Flash-Next by Qwen) is the 125B-A6B preview of the Qwen4 architecture: a hybrid GatedDeltaNet mixture of experts with Qwen Sparse Attention and a 51B-parameter n-gram memory. MTPLX runs its native multi-token-prediction head as an exact speculative decoder. On MTPLX 2.11.3 a MacBook Pro M5 Max with 128 GB decoded an OpenCode request at 125.8 tok/s, a 9k-token code prompt at 79.3 tok/s (62.5 on 2.11.2), a 109k-token OpenCode turn at 61.8 tok/s and a 200k-token turn at 50.3 tok/s, with the output the model's own at any temperature. MTPLX 2.10.0 (29 August 2026) was the first Apple Silicon backend for the family. On an M5 Max with 128 GB, MTPLX 2.12.0 (23 September 2026) reads a 4,061-token prompt 85 percent faster than 2.11.3, with the first token after 2.88 s instead of 5.26 s. Three packs: Bare Speed for Macs with 96 GB or more, Optimized Speed for 128 GB or more, and the new 8-bit Optimized Quality for 256 GB or more.

Three packs

All three packs keep the model's MTP head and run the same exact speculative path. They differ in quantization. Optimized Speed keeps the Qwen Sparse Attention projections at 8-bit, so the attention pathway that steers long contexts keeps its precision; it is the recommended build. Bare Speed puts every expert at flat 4-bit for the quickest Flash-Next speeds. Optimized Quality, new in 2.12.0, is the 8-bit build for Macs with 256 GB or more (details below).

PackQuantDownloadResident weightsPick it for
Optimized SpeedDynamic 4-bit, 64-weight groups; Qwen Sparse Attention projections at 8-bit115.1 GB~83 GB + working setRecommended. Higher quality, slightly slower.
Bare SpeedFlat 4-bit, 64-weight groups, nothing promoted106.3 GB~74 GB + working setQuickest Flash-Next speeds for chat and coding.
Optimized Quality8-bit main model and draft head, 64-weight groups; structural weights in BF16; n-gram table at 4-bit, 32-weight groups169.96 GB~128.5 GiBMacs with 256 GB or more. Its speed has not been measured yet.

All three downloads include the 32 GB n-gram embedding table. It ships as a separate sidecar file that MTPLX streams from SSD on every Mac since 2.12.0, so the weights stay resident and the table does not have to. All three packs keep the vision tower. Context window: 262,144 tokens. The base model is Qwen/Qwen3.8-Flash-Next under the Qwen Community License; the upstream model card is preserved in each repo as README-upstream-qwen.md.

Qwen 3.8 Flash Next 8-bit: Optimized Quality

New in MTPLX 2.12.0, Optimized Quality is an 8-bit build of Qwen 3.8 Flash Next for Macs with 256 GB or more, such as a Mac Studio with 256 GB or 512 GB. The main model and the draft head are at 8 bits with 64-weight groups, twice the precision of Optimized Speed. The structural weights stay in BF16, and the n-gram table is at 4 bits with 32-weight groups. The download is 169.96 GB, including the 32 GB n-gram table. With the table streaming from SSD, the weights need about 128.5 GiB of memory, which leaves about 59.5 GiB for context and the session cache on a 256 GB Mac. All 2,860 tensors were checked, and sampled values in 25 groups were compared with the BF16 source.

It has not been run on a 256 GB Mac yet, so on those Macs the app and the CLI recommend Optimized Speed first and list Optimized Quality second. Its speed has not been measured. The 8-bit weights move twice the bytes per token of Optimized Speed, so it decodes slower. Until it has run on a 256 GB Mac, mtplx inspect and mtplx serve label it "Official MTPLX pack, qualification pending"; the label does not stop it from running. The pack needs MTPLX 2.12.0 or later and carries the Flash-Next sampling settings (temperature 1.0, top-p 0.95, top-k 20). In the app, pick "Qwen 3.8 Flash-Next Optimized Quality". From the command line:

mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality

Build it yourself. Forge builds this pack with the flash-next-optimized-quality recipe. One command builds the pack, loads it, and tests chat, a tool call and an image:

bash scripts/build_flash_next_quality_pack.sh Qwen/Qwen3.8-Flash-Next <output directory>

Forge converts Flash-Next source models since 2.12.0 (#508 by @bpmforge). It used to stop with Model type qwen4_exp not supported. The 51B n-gram table is now quantized one shard at a time straight into its own file, and the rest of the model is converted one tensor at a time with MTPLX's own model code. A follow-up reads the converted files back through the loaders the runtime uses.

Faster prompts in 2.12.0

MTPLX 2.12.0 against 2.11.3 on an M5 Max with 128 GB, measured on 22 September 2026. Both versions used the same Python runtime and were booted in the order 2.11.3, 2.12.0, 2.12.0, 2.11.3, with one model loaded at a time and the fans at maximum. Thinking was off, each answer was 512 tokens, sampling used the model's own settings with seed 1731, and the GPU memory limit was raised to 120 GiB (the default is about 96 GiB). From the 2.12.0 release notes.

Flash-NextMTPLX 2.11.3MTPLX 2.12.0
Time to the first token, 4,061-token prompt5.26 s2.88 s
Prompt processing, 4,061-token prompt786 tok/s1,453 tok/s (+85%)
Decode after a 4,061-token prompt74.4 tok/s74.1 tok/s (the same)
Time to the first token, 65,502-token prompt85.5 s60.1 s
Prompt processing, 65,502-token prompt768 tok/s1,094 tok/s (+42%)
Decode after a 65,502-token prompt56.1 tok/s63.5 tok/s (+13%)

At 65,502 tokens, 2.11.3 copied the whole key and value cache on every verify round. 2.12.0 sizes the cache so that every write happens in place, with identical output. On M5 chips, Flash-Next now reads prompts 4,096 tokens at a time and switches to sparse attention from 16K tokens; other chips keep 2,048-token chunks and the 32K switch point.

Five prompt steps that now run as single GPU kernels were added after that run, so the release is faster than the table shows. In their own paired run on the same M5 Max, alternating boots without and with the kernels, they added 16 percent to prompt processing at 16,376 tokens and 9 percent at 65,529 tokens. Peak memory was the same (92.0 GB at 64K), and every position's logits and 256 greedy tokens matched on code and prose at 16K and on code at 64K.

PromptWithout the kernelsWith the kernelsTime to the first token
16,376 tokens1,459 tok/s1,695 tok/s (+16%)11.3 s to 9.7 s
65,529 tokens1,415 tok/s1,548 tok/s (+9%)46.6 s to 42.6 s

Four of these kernels run on every Mac, and each Flash-Next load checks them on that Mac's GPU first. The wide input projection step runs only on M5 chips. M1 to M4 keep the quantized multiply and 2,048-token chunks, and their speed was not measured.

Measured speeds

MTPLX 2.11.3 (17 September 2026), MacBook Pro M5 Max with 128 GB, fans verified at maximum, the Optimized Speed pack, sampled at the model's own settings (temperature 1.0, top-p 0.95, top-k 20). Every row is from the 2.11.3 release notes.

Runtok/sConditions
One OpenCode request125.81,301 tokens generated, 18,539-token prompt with 18,364 tokens served from cache, MTP depth 3, OpenCode Desktop, 16 September 2026. The fastest measured request.
OpenCode coding tasks, per model request54.6 to 122.4real multi-file tasks on the installed daemon; seven of eight requests served from the session cache
9k-token code prompt79.31,500 tokens generated, seeded sampler, thinking off, two alternating boots each; 62.5 on 2.11.2 on the same Mac and runtime (+27 percent)
109k-token OpenCode turn61.8mean of two runs, 60.8 and 62.7; 48.8 on the same build before the launcher and depth-policy fix, memory unchanged
200k-token OpenCode turn, warm50.3mean of two runs, 50.0 and 50.6; the cold 200,073-token prompt in the installed app decoded at 42.5 and its warm follow-up at 50.6
Full 45k to 56k-token generations66.8whole turn, the Flappy Bird prompt at effort xhigh; 66.5 and 62.6 on 2.11.2 on alternating boots
Rewrite after a 4 ms restore in the app73.2a 54,530-token conversation restored from the session cache in 4 ms, then a 42,200-token rewrite; a 96,760-token conversation restored in 8 ms

Exactness on the same release: at temperature 1, top-p 0.95, top-k 20, a thousand four-token draws from the fast path match a thousand from the plain path within the plain path's own noise at every joint length. 261,120-token prompts decode. Eight exactness defects found in MTPLX's own engine were fixed in 2.11.3, each with a test.

Decode by context length, MTPLX 2.11 (4 September 2026) against 2.10.2 on the same M5 Max with 128 GB: same-hour pairs, temperature 1, copy lane on in both arms, coding prompts, from the 2.11 release note.

ContextMTPLX 2.10.2MTPLX 2.11Notes
16k tokens, 1,024-token answer53.2 tok/s68.4 tok/s+29%. Round 48.2 to 39.2 ms; cold-prefill time to first token 15.1 to 14.3 s; peak memory 89.4 GB, flat.
100k tokens47.5 tok/s60.9 tok/s+28%. Round 54.5 to 43.4 ms; cold-prefill time to first token 117.0 to 113.2 s; peak 97.9 GB against 98.1 GB.
206k tokens32.2 tok/s44.2 tok/s+37%. At this rung the compiled verify lane is memory-gated on 128 GB, so 44.2 is the shipped default there.

The lane behind those rows is a compiled verify path plus a set of decode and prefill items adapted from PR #391 by @davidtai, ported commit by commit under his name, together with MTPLX's own rows-gather fence, pack-checked kernel rules, per-request memory gate and n-gram table pre-read that keep the lane exact and on by default. Every item is token-identical to the tree before it at temperature 0, and every knob honors an explicit export, =0 included.

Warm agent turns: the release's agent-session gate at 40k tokens of context runs 0.15 s to first token, 0.01 s of dead time and 66 to 84 tok/s (a 43k-token OpenCode session replayed against the live daemon). Before 2.11 a tool turn after a file write re-prefilled 3,535 tokens in 3.7 s; it re-prefills 20 tokens in 0.12 s. First token after a pause of 3 to 90 s: 0.08 s server time, down from 1.06 s, because the engine keeps its 77 GiB working set GPU-resident while a request has completed in the last ten minutes.

The two pack cards, measured at launch through the MTPLX server (mtplx serve) on an M5 Max, fans verified at max, single stream, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20), sampled output, on the Flash-Next backend that shipped in MTPLX 2.10.0 on 29 August 2026. Same coding task for both packs.

RunOptimized SpeedBare Speed
Coding task, MTP speculative decode (the default)73.5 tok/s75.9 tok/s
Same task, plain autoregressive43.8 tok/s47.0 tok/s
Speculative multiplier through the product serve path1.7x1.6x

The 2.10.1 release note (30 August 2026) added block-sparse prefill for Flash-Next: peak memory on a 98k-token prompt fell from 91.4 to 83.0 GB, and a 262,144-token cold prompt completed at 87.4 GB peak, where 2.10.0 needed 119 GB or did not complete. On the macOS 27 betas, 2.10.x answered prompts past the 32k sparse-prefill crossover with HTTP 500; 2.11 fixes the seven kernel sites, and the reporters' receipts on the beta read 64.4 s for a 70k-token prompt and 84.9 s for 95k.

Other engines

The one table that runs MTPLX and mlx-serve on one machine is mlx-serve's own, in its repository at tag v26.9.3 (docs/mtp-acceptance-port.md, lines 42 to 46): on an M5 Max at temperature 1, top-p 0.95, thinking off, MTP depth 3, MTPLX exact decoding measured 102.0 tok/s on a short prompt and 90.9 tok/s at about 16K tokens, while mlx-serve's exact default measured 92.5 and 82.4. The only rows where mlx-serve is ahead are the two opt-in modes its documentation calls not distribution-exact. That is their measurement of an MTPLX build on their machine. Their published release table is an M4 Max at temperature 0: 80 tok/s for their Flash Next pack on 26.9.3. Pack quality is their number too: their 4-bit Flash Next pack agrees with bf16 on 85.6 percent of greedy tokens. The full comparison, with file and line for every quote: MTPLX vs mlx-serve. Other engines: oMLX, LM Studio, Ollama, llama.cpp.

RAM and Macs

  • 96 GB: Bare Speed. Since 2.12.0 the app and the CLI point 96 GB Macs to Bare Speed, which peaks at 78 GiB, while Optimized Speed peaks at 87 GiB. On 96 GB, the catalog, the memory planner and the verify memory check share one 84 GiB budget for Flash-Next, with a planned window of 86,016 tokens on M5 and 20,480 tokens on chips without tensor units. The 2.10.1 release confirmed that 96 GB Macs load Flash-Next.
  • 128 GB or more: Optimized Speed, the recommended build, or Bare Speed, with the n-gram table streaming from SSD. Resident weights about 83 GB (Optimized Speed) or about 74 GB (Bare Speed) plus working set.
  • 256 GB or more: Optimized Speed is the first recommendation and Optimized Quality the second, because the 8-bit pack has not run on a 256 GB Mac yet. On M3, M4 and M5 Macs the app and the CLI also list Optimized Quality as a Flash-Next option from 192 GB.
  • Context window: 262,144 tokens. The memory governor (2.10.0) prints engine budget, weights, resolved context window and session bank in the serve banner; requests that cannot fit are refused up front with HTTP 507 (2.10.2, 1 September 2026). Since 2.11 an explicit MTPLX_MEMORY_LIMIT_BYTES is the engine budget both ways, so a 96 GB Mac serving the Bare Speed pack under a limit of 80G runs instead of being refused, and mtplx serve --allow-swap (or the app's Settings > Memory card) admits prompts past the memory fit for operators who accept swap.
  • 128 GB Macs: the compiled verify lane promotes every sparse-attention layer's state into padded banks, about 28 KB per context token, so a per-request memory gate hands roughly the last 20 percent of the context window to the plain verify instead. The engine envelope stays at three quarters of the machine.
  • Image input: PNG, JPEG and WebP since 2.10.1 (30 August 2026).
  • Smaller Macs: Qwen 3.8 27B on 32 GB or more. On 16 to 31 GB, the app suggests Ternary Bonsai 2 27B first and MiMo V2.6 Qwen 9B second on M3, M4 and M5 Macs, and Qwen 3.5 9B and 4B fit too. The app checks your Mac before recommending anything.

Install

Mac app: download the DMG, pick "Qwen 3.8 Flash-Next Optimized Speed", "Qwen 3.8 Flash-Next Bare Speed" or, on Macs with 256 GB or more, "Qwen 3.8 Flash-Next Optimized Quality". The app downloads the pack, sets up its engine, and measures your machine to pick the fastest decoding depth.

Command line:

brew install youssofal/mtplx/mtplx
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed

For Bare Speed, pass --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Bare-Speed; the Optimized Quality command is in its section above. Then point OpenCode, Pi, Claude Code, Cline, Cursor or anything that speaks the OpenAI or Anthropic API at http://127.0.0.1:8000. The served model ids are mtplx-flash-next-optimized-speed and mtplx-flash-next-bare-speed; mtplx connect claude-code and mtplx connect opencode print the exact client config. Setup pages for each client are in the docs. The serving contract ships inside mtplx_runtime.json; MTPLX reads it on load. The Turbo profile is the default for the Flash-Next packs. Concurrent requests are served one at a time on Flash-Next (2.11, #420: the family's sparse-attention cache has no batch merge yet, and the two requests that used to raise a 500 now queue); the 27B keeps batching.

How it is built

Optimized Speed:

  • MoE experts and dense matrices at 4-bit with 64-weight groups; the Qwen Sparse Attention projections promoted to 8-bit, the quality edge over Bare Speed.
  • The GDN convolution and recurrent-state parameters, every norm, the QSA indexer, and the MTP head stay 16-bit.
  • The n-gram embedding table ships as a separate ngram-table.safetensors sidecar that MTPLX streams from SSD. The vision tower is preserved in the weights.

Bare Speed is every MoE expert and dense matrix at 4-bit with 64-weight groups, nothing promoted, with the same 16-bit set and the same n-gram sidecar. Optimized Quality puts the main model and the draft head at 8 bits with 64-weight groups, keeps the structural weights in BF16, and stores the n-gram table at 4 bits with 32-weight groups. All three packs carry their sampling contract (temperature 1.0, top-p 0.95, top-k 20, the official Qwen 3.8 contract) in mtplx_runtime.json.

Exactness

Speculation in MTPLX is exact. Drafts from the MTP head are accepted with the probability-ratio rule and rejected drafts are resampled from the residual, so what you sample is what the model would have sampled without speculation, at any temperature. The draft sampler is a speed knob only. The quantization is the one approximation.