Home›Models›granite-4.2-8b

Model sheet

granite-4.2-8b

Architecture dense40 LayersLicense apache-2.0Maximum context 128kVocabulary 100,352

At Q4_K_M this is a 4.98 GiB file and it fits on 18 of the 18 machine profiles we measure. What it costs to run is not the file: it is the context — and the same content in German costs ×1.61 what it costs in English.

At a glance
Size at Q4_K_M4.98 GiB
Context memory at 8k1.25 GiB
Fits on18 / 18
German cost×1.61
Smallest machine that takes itPC with an RTX 4060 8 GB
4.98 GiB

Model size at Q4_K_M

Measured by Local AI Scope

1.25 GiB

Context memory at 8k tokens

Measured by Local AI Scope

×1.61

What German costs versus English

Measured by Local AI Scope

Sizes

Sizes of the real published files

These are the sizes of the files actually published by the author, added up when the model ships split into parts. Not an estimate from the parameter count.

Size of the published file for each compression
Compression Model size
Q8_0 8.70 GiB
Q6_K 6.72 GiB
Q5_K_M 5.82 GiB
Q5_K_S 5.68 GiB
Q4_K_M 4.98 GiB
Q4_K_S 4.74 GiB
Q3_K_M 4.05 GiB
Context

Context memory

Computed from the attention pattern this model declares layer by layer, not from the generic formula. That is why the figure is often far smaller than other calculators tell you. This model’s config.json does not state the size of each attention head, and that number is what the context memory is made of. We derive it from two figures it does state, 4,096 hidden units over 32 attention heads, which is what the size of a head means and what the reference library assumes by default: 128. Every figure in this section rests on that derivation, so it is labelled here rather than left implicit.

  • 8k tokens → 1.25 GiB
  • 32k tokens → 5.00 GiB
  • 128k tokens → 20.00 GiB
Languages

What it costs in each language

Measured by running this model’s own tokenizer over the same content in four languages. Tokens are what you pay for in context, in memory and in your cloud bill.

Token cost per kind of text and language, with English as the baseline
Kind of text EN ES FR DE
Code — ×1.22 ×1.42 ×1.60
Contracts — ×1.39 ×1.47 ×1.76
Business email — ×1.27 ×1.41 ×1.48
Support tickets — ×1.39 ×1.63 ×1.55

English is the baseline: every figure is how many times more tokens the same content costs in that language. Average across the whole corpus: DE ×1.61.

Machines

Where it fits

Fits means the model, its context memory and the working headroom all fit at the stated context, leaving the system its share. 18 of the 18 machine profiles tested (Q4_K_M, 8k tokens).

The 18 machine profiles, with the memory each one leaves and the largest context this model holds on it
Machine Fits Memory left Largest context that fits
PC with an RTX 4060 8 GB yes 7.20 GiB 8k
PC with an RTX 3060 12 GB yes 11.20 GiB 32k
PC with an RTX 4070 12 GB yes 11.20 GiB 32k
Mac with M4 and 16 GB unified memory yes 13.00 GiB 32k
Office laptop, CPU only, 16 GB yes 13.00 GiB 32k
Mac mini with M6 and 16 GB yes 13.00 GiB 32k
Mac with M4 Pro and 24 GB yes 21.00 GiB 64k
Workstation with an RTX 4090 24 GB yes 23.20 GiB 64k
Mini-PC with an NPU and 32 GB LPDDR5X yes 29.00 GiB 128k
Mac mini with M6 and 32 GB yes 29.00 GiB 128k
Server with 2× RTX 3090 (48 GB) yes 47.20 GiB 128k
Mac with M4 Max and 64 GB yes 61.00 GiB 128k
Mac mini with M5 Pro and 64 GB yes 61.00 GiB 128k
Xiaomi AI Cube, 80 GB — engineering prototype yes 77.00 GiB 128k
PC with Ryzen AI Max+ 395 and 128 GB unified yes 96.00 GiB 128k
Mac Studio with M5 Max and 128 GB yes 125.00 GiB 128k
PC with Ryzen AI Max+ PRO 495 and 192 GB unified yes 160.00 GiB 128k
Mac Studio with M5 Ultra and 512 GB yes 509.00 GiB 128k