qwen3.6-35b-a3b
Architecture mixture of experts40 LayersLicense apache-2.0Maximum context 256kVocabulary 248,320
At Q4_K_M this is a 20.61 GiB file and it fits on 4 of the 10 machine profiles we measure. What it costs to run is not the file: it is the context — and the same content in German costs ×1.31 what it costs in English.
Model size at Q4_K_M
Measured by Local AI Scope
Context memory at 8k tokens
Measured by Local AI Scope
What German costs versus English
Measured by Local AI Scope
Sizes of the real published files
These are the sizes of the files actually published by the author, added up when the model ships split into parts. Not an estimate from the parameter count.
| Compression | Model size |
|---|---|
Q8_0 |
34.37 GiB |
Q6_K |
27.30 GiB |
Q5_K_M |
24.64 GiB |
Q5_K_S |
23.23 GiB |
Q4_K_M |
20.61 GiB |
Q4_K_S |
19.46 GiB |
IQ4_XS |
16.51 GiB |
Q3_K_M |
15.46 GiB |
Context memory
Computed from the attention pattern this model declares layer by layer, not from the generic formula. That is why the figure is often far smaller than other calculators tell you.
- 8k tokens → 0.16 GiB
- 32k tokens → 0.62 GiB
- 128k tokens → 2.50 GiB
What it costs in each language
Measured by running this model’s own tokenizer over the same content in four languages. Tokens are what you pay for in context, in memory and in your cloud bill.
| Kind of text | EN | ES | FR | DE |
|---|---|---|---|---|
| Code | — | ×1.20 | ×1.30 | ×1.46 |
| Contracts | — | ×1.16 | ×1.23 | ×1.31 |
| Business email | — | ×1.12 | ×1.22 | ×1.17 |
| Support tickets | — | ×1.24 | ×1.40 | ×1.25 |
English is the baseline: every figure is how many times more tokens the same content costs in that language. Average across the whole corpus: DE ×1.31.
Where it fits
Fits means the model, its context memory and the working headroom all fit at the stated context, leaving the system its share. 4 of the 10 machine profiles tested (Q4_K_M, 8k tokens).
| Machine | Fits | Memory left | Largest context that fits |
|---|---|---|---|
| PC with an RTX 4060 8 GB | no | 7.20 GiB | — |
| PC with an RTX 3060 12 GB | no | 11.20 GiB | — |
| PC with an RTX 4070 12 GB | no | 11.20 GiB | — |
| Mac with M4 and 16 GB unified memory | no | 13.00 GiB | — |
| Office laptop, CPU only, 16 GB | no | 13.00 GiB | — |
| Mac with M4 Pro and 24 GB | no | 21.00 GiB | — |
| Workstation with an RTX 4090 24 GB | yes | 23.20 GiB | 32k |
| Mini-PC with an NPU and 32 GB LPDDR5X | yes | 29.00 GiB | 128k |
| Server with 2× RTX 3090 (48 GB) | yes | 47.20 GiB | 256k |
| Mac with M4 Max and 64 GB | yes | 61.00 GiB | 256k |