Home›Hardware›PC with an RTX 4060 8 GB
PC with an RTX 4060 8 GB
NVIDIA graphics card8 GiB272 GB/s15.0 TFLOPS
This machine leaves 7.20 GiB for the model once the system has its share, and it takes 5 of the 18 models we measure at the reference compression. What decides how fast it answers is not the chip: it is the 272 GB/s of memory bandwidth.
Left for the model
Estimated
Memory bandwidth
Estimated
Models that fit
Measured by Local AI Scope
The memory you can actually use
This is why the number on the box is not the number you get. The desktop, the compositor and the graphics context never give their share back, and a model that does not leave them room does not load — it crashes.
| Memory | GiB |
|---|---|
| Memory on the box | 8 |
| Reserved for the system | −0.80 |
| Left for the model | 7.20 |
Memory bandwidth
Two different limits, and people confuse them constantly. While the model is writing its answer, bandwidth rules: every single token means reading the weights again, so tokens per second is roughly bandwidth divided by model size. While the model is reading your document, compute rules: the whole input is processed at once. That is why a laptop with no GPU can take minutes before the first word appears and then type at a tolerable pace.
272 GB/s — Derived, not published by the manufacturer · the manufacturer publishes the bus width but neither the memory speed nor the bandwidth. This figure is the usual arithmetic derivation (17 Gbps x 128 bit / 8 = 272 GB/s) and we could not confirm it at source (Source).
Compute: 15.0 TFLOPS (FP32 shader TFLOPS) — Vendor declared, under a different name · the manufacturer publishes this number as FP32 shader TFLOPS, not as FP16; checked against the manufacturer’s page on 2026-08-25 (Source). The figures in this line do not all come from the same measure across machines, because each manufacturer publishes a different one. Compare them between machines only with that in mind.
Which models fit here, and which do not
Fitting means three things at once fit in the memory left over: the model file, its context memory, and the runtime’s working headroom. Context memory is computed from each model’s declared attention pattern layer by layer, not from the generic formula — which is why several models fit here that other calculators say do not.
| Model | Compression shown | Model size | 8k | 32k | 128k | Largest context that fits |
|---|---|---|---|---|---|---|
| qwen3-1.7b | Q4_K_M |
1.03 GiB | yes | yes | — | 32k |
| qwen3-4b-2507 | Q4_K_M |
2.33 GiB | yes | no | no | 16k |
| gemma4-e2b | Q4_K_M |
2.89 GiB | yes | yes | yes | 128k |
| gemma4-e4b | Q4_K_M |
4.63 GiB | yes | yes | no | 32k |
| qwen3-8b | Q4_K_M |
4.68 GiB | yes | no | — | 8k |
| gemma4-12b | Q4_K_M |
6.63 GiB | no | no | no | — |
| gpt-oss-20b | Q4_K_M |
10.83 GiB | no | no | no | — |
| mistral-small-24b | Q4_K_M |
13.35 GiB | no | no | no | — |
| qwen3.8-27b | Q4_K_M |
15.33 GiB | no | no | no | — |
| qwen3.6-27b | Q4_K_M |
15.66 GiB | no | no | no | — |
| gemma4-26b-a4b | Q4_K_M |
15.78 GiB | no | no | no | — |
| gemma4-31b | Q4_K_M |
17.07 GiB | no | no | no | — |
| qwen3-coder-30b | Q4_K_M |
17.28 GiB | no | no | no | — |
| qwen3.6-35b-a3b | Q4_K_M |
20.61 GiB | no | no | no | — |
| gpt-oss-120b | Q4_K_M |
58.46 GiB | no | no | no | — |
| llama4-scout-17b | Q4_K_M |
60.87 GiB | no | no | no | — |
| deepseek-v4-flash | IQ4_XS |
127.28 GiB | no | no | no | — |
| glm-5.2 | Q4_K_M |
433.83 GiB | no | no | no | — |
Sizes are the real published files at the reference compression, one step per row: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Every row states which one it is showing. Q4_K_M is our quality floor: below it the loss is audible in the answers. A model that would only fit here at a harsher compression is listed as not fitting, on purpose. A dash in a context column means the model itself does not offer that context, so there is nothing to fit.
Models that fit — 5 of the 18 models measured
- qwen3-1.7b
Q4_K_M— context 32k tokens - qwen3-4b-2507
Q4_K_M— context 16k tokens - gemma4-e2b
Q4_K_M— context 128k tokens - gemma4-e4b
Q4_K_M— context 32k tokens - qwen3-8b
Q4_K_M— context 8k tokens
Models that do not fit — 13 of the 18 models measured
- gemma4-12b
Q4_K_M— short by 0.67 GiB - gpt-oss-20b
Q4_K_M— short by 4.40 GiB - mistral-small-24b
Q4_K_M— short by 7.45 GiB - qwen3.8-27b
Q4_K_M— short by 9.06 GiB - qwen3.6-27b
Q4_K_M— short by 9.39 GiB - gemma4-26b-a4b
Q4_K_M— short by 9.61 GiB - gemma4-31b
Q4_K_M— short by 11.95 GiB - qwen3-coder-30b
Q4_K_M— short by 11.13 GiB - qwen3.6-35b-a3b
Q4_K_M— short by 14.17 GiB - gpt-oss-120b
Q4_K_M— short by 52.08 GiB - llama4-scout-17b
Q4_K_M— short by 55.10 GiB - deepseek-v4-flash
IQ4_XS— short by 121.09 GiB - glm-5.2
Q4_K_M— short by 441.93 GiB