Home›Models›qwen3.8-flash-next

Model sheet

qwen3.8-flash-next

Architecture mixture of experts48 LayersLicense qwen-community-1.0Maximum context 256kVocabulary 248,320

At IQ4_XS this is an 87.25 GiB file and it fits on 4 of the 18 machine profiles we measure. Here the file is what fills the memory: even at its maximum context of 256k tokens, the context adds 6.00 GiB — and the same content in German costs ×1.31 what it costs in English.

At a glance
Size at IQ4_XS87.25 GiB
Context memory at 8k0.19 GiB
Fits on4 / 18
German cost×1.31
Smallest machine that takes itPC with Ryzen AI Max+ 395 and 128 GB unified
87.25 GiB

Model size at IQ4_XS

Measured by Local AI Scope

0.19 GiB

Context memory at 8k tokens

Measured by Local AI Scope

×1.31

What German costs versus English

Measured by Local AI Scope

Sizes

Sizes of the real published files

These are the sizes of the files actually published by the author, added up when the model ships split into parts. Not an estimate from the parameter count.

Size of the published file for each compression
Compression Model size
Q8_0 175.30 GiB
IQ4_XS 87.25 GiB

This model also reads images. For that it needs a second file, the vision projector (mmproj-F16.gguf, 0.84 GiB), published next to the model and not included in the figures on this page: they count text only. If you are going to send it images, add it to the model size.

Context

Context memory

Computed from the attention pattern this model declares layer by layer, not from the generic formula. That is why the figure is often far smaller than other calculators tell you. Besides that memory, this model keeps a second, smaller cache for its sparse-attention indexer: about 3 KiB per token, 0.02 GiB at 8k and 0.75 GiB at its maximum context. It is derived from the code and the file header, not declared by the author, and it is not added to the figures on this page. Added in, it changes none of the verdicts in the machine table.

  • 8k tokens → 0.19 GiB
  • 32k tokens → 0.75 GiB
  • 128k tokens → 3.00 GiB
Languages

What it costs in each language

Measured by running this model’s own tokenizer over the same content in four languages. Tokens are what you pay for in context, in memory and in your cloud bill.

Token cost per kind of text and language, with English as the baseline
Kind of text EN ES FR DE
Code — ×1.20 ×1.30 ×1.46
Contracts — ×1.16 ×1.23 ×1.31
Business email — ×1.12 ×1.22 ×1.17
Support tickets — ×1.24 ×1.40 ×1.25

English is the baseline: every figure is how many times more tokens the same content costs in that language. Average across the whole corpus: DE ×1.31.

Machines

Where it fits

Fits means the model, its context memory and the working headroom all fit at the stated context, leaving the system its share. 4 of the 18 machine profiles tested (IQ4_XS, 8k tokens).

The 18 machine profiles, with the memory each one leaves and the largest context this model holds on it
Machine Fits Memory left Largest context that fits
PC with an RTX 4060 8 GB no 7.20 GiB —
PC with an RTX 3060 12 GB no 11.20 GiB —
PC with an RTX 4070 12 GB no 11.20 GiB —
Mac with M4 and 16 GB unified memory no 13.00 GiB —
Office laptop, CPU only, 16 GB no 13.00 GiB —
Mac mini with M6 and 16 GB no 13.00 GiB —
Mac with M4 Pro and 24 GB no 21.00 GiB —
Workstation with an RTX 4090 24 GB no 23.20 GiB —
Mini-PC with an NPU and 32 GB LPDDR5X no 29.00 GiB —
Mac mini with M6 and 32 GB no 29.00 GiB —
Server with 2× RTX 3090 (48 GB) no 47.20 GiB —
Mac with M4 Max and 64 GB no 61.00 GiB —
Mac mini with M5 Pro and 64 GB no 61.00 GiB —
Xiaomi AI Cube, 80 GB — engineering prototype no 77.00 GiB —
PC with Ryzen AI Max+ 395 and 128 GB unified yes 96.00 GiB 128k
Mac Studio with M5 Max and 128 GB yes 125.00 GiB 256k
PC with Ryzen AI Max+ PRO 495 and 192 GB unified yes 160.00 GiB 256k
Mac Studio with M5 Ultra and 512 GB yes 509.00 GiB 256k

One piece of this model behaves differently depending on the program that runs it. 26.82 GiB of the IQ4_XS file are an n-gram table (per_layer_token_embd) that is only ever read a few rows at a time. llama.cpp v0.5.0 and later leave it on disk by default and read those rows on demand, so what has to fit in memory drops to 60.43 GiB. On that basis it would also fit on Xiaomi AI Cube, 80 GB — engineering prototype, up to 256k tokens of context. Closest miss: Mac with M4 Max and 64 GB · Mac mini with M5 Pro and 64 GB, short by 0.20 GiB. The table above counts the whole file, because that holds for every program; the saving is checked only for llama.cpp v0.5.0, in its code and in the file header. That same code marks its support for this architecture as provisional, pending a full reimplementation (read on 29 September 2026).