HomeModelsglm-5.2

Model sheet

glm-5.2

Architecture dense78 LayersLicense mitMaximum context 1,024kVocabulary 154,880

At Q4_K_M this is a 433.83 GiB file and it fits on none of the 10 machine profiles we measure at 8k tokens. The file is not what fills your memory — the context is.

At a glance
Size at Q4_K_M433.83 GiB
Context memory at 8k29.25 GiB
Fits on0 / 10
German cost×1.49
433.83 GiB

Model size at Q4_K_M

Measured by Local AI Scope

29.25 GiB

Context memory at 8k tokens

Measured by Local AI Scope

×1.49

What German costs versus English

Measured by Local AI Scope

Sizes

Sizes of the real published files

These are the sizes of the files actually published by the author, added up when the model ships split into parts. Not an estimate from the parameter count.

Size of the published file for each compression
Compression Model size
Q8_0 746.32 GiB
Q6_K 582.88 GiB
Q5_K_M 522.31 GiB
Q5_K_S 491.05 GiB
Q4_K_M 433.83 GiB
Q4_K_S 406.46 GiB
IQ4_XS 340.22 GiB
Q3_K_M 319.20 GiB
Context

Context memory

Computed from the attention pattern this model declares layer by layer, not from the generic formula. That is why the figure is often far smaller than other calculators tell you.

  • 8k tokens → 29.25 GiB
  • 32k tokens → 117.00 GiB
  • 128k tokens → 468.00 GiB
Languages

What it costs in each language

Measured by running this model’s own tokenizer over the same content in four languages. Tokens are what you pay for in context, in memory and in your cloud bill.

Token cost per kind of text and language, with English as the baseline
Kind of text EN ES FR DE
Code ×1.26 ×1.41 ×1.64
Contracts ×1.29 ×1.36 ×1.52
Business email ×1.22 ×1.32 ×1.32
Support tickets ×1.33 ×1.51 ×1.41

English is the baseline: every figure is how many times more tokens the same content costs in that language. Average across the whole corpus: DE ×1.49.

Machines

Where it fits

Fits means the model, its context memory and the working headroom all fit at the stated context, leaving the system its share. 0 of the 10 machine profiles tested (Q4_K_M, 8k tokens).

The ten machine profiles, with the memory each one leaves and the largest context this model holds on it
Machine Fits Memory left Largest context that fits
PC with an RTX 4060 8 GB no 7.20 GiB
PC with an RTX 3060 12 GB no 11.20 GiB
PC with an RTX 4070 12 GB no 11.20 GiB
Mac with M4 and 16 GB unified memory no 13.00 GiB
Office laptop, CPU only, 16 GB no 13.00 GiB
Mac with M4 Pro and 24 GB no 21.00 GiB
Workstation with an RTX 4090 24 GB no 23.20 GiB
Mini-PC with an NPU and 32 GB LPDDR5X no 29.00 GiB
Server with 2× RTX 3090 (48 GB) no 47.20 GiB
Mac with M4 Max and 64 GB no 61.00 GiB

None of the machine profiles we measure takes this model at the reference compression, not even at the smallest context.