granite-4.2-8b
Architecture dense40 LayersLicense apache-2.0Maximum context 128kVocabulary 100,352
At Q4_K_M this is a 4.98 GiB file and it fits on 18 of the 18 machine profiles we measure. What it costs to run is not the file: it is the context — and the same content in German costs ×1.61 what it costs in English.
Model size at Q4_K_M
Measured by Local AI Scope
Context memory at 8k tokens
Measured by Local AI Scope
What German costs versus English
Measured by Local AI Scope
Sizes of the real published files
These are the sizes of the files actually published by the author, added up when the model ships split into parts. Not an estimate from the parameter count.
| Compression | Model size |
|---|---|
Q8_0 |
8.70 GiB |
Q6_K |
6.72 GiB |
Q5_K_M |
5.82 GiB |
Q5_K_S |
5.68 GiB |
Q4_K_M |
4.98 GiB |
Q4_K_S |
4.74 GiB |
Q3_K_M |
4.05 GiB |
Context memory
Computed from the attention pattern this model declares layer by layer, not from the generic formula. That is why the figure is often far smaller than other calculators tell you. This model’s config.json does not state the size of each attention head, and that number is what the context memory is made of. We derive it from two figures it does state, 4,096 hidden units over 32 attention heads, which is what the size of a head means and what the reference library assumes by default: 128. Every figure in this section rests on that derivation, so it is labelled here rather than left implicit.
- 8k tokens → 1.25 GiB
- 32k tokens → 5.00 GiB
- 128k tokens → 20.00 GiB
What it costs in each language
Measured by running this model’s own tokenizer over the same content in four languages. Tokens are what you pay for in context, in memory and in your cloud bill.
| Kind of text | EN | ES | FR | DE |
|---|---|---|---|---|
| Code | — | ×1.22 | ×1.42 | ×1.60 |
| Contracts | — | ×1.39 | ×1.47 | ×1.76 |
| Business email | — | ×1.27 | ×1.41 | ×1.48 |
| Support tickets | — | ×1.39 | ×1.63 | ×1.55 |
English is the baseline: every figure is how many times more tokens the same content costs in that language. Average across the whole corpus: DE ×1.61.
Where it fits
Fits means the model, its context memory and the working headroom all fit at the stated context, leaving the system its share. 18 of the 18 machine profiles tested (Q4_K_M, 8k tokens).
