HomeHardwareWorkstation with an RTX 4090 24 GB

Hardware sheet

Workstation with an RTX 4090 24 GB

NVIDIA graphics card24 GiB1,008 GB/s82.6 TFLOPS

This machine leaves 23.20 GiB for the model once the system has its share, and it takes 14 of the 18 models we measure at the reference compression. What decides how fast it answers is not the chip: it is the 1,008 GB/s of memory bandwidth.

At a glance
Left for the model23.20 GiB
Memory bandwidth1,008 GB/s
Models that fit14 / 18
Largest model it takesqwen3.6-35b-a3b
23.20 GiB

Left for the model

Estimated

1,008 GB/s

Memory bandwidth

Vendor declared

14 / 18

Models that fit

Measured by Local AI Scope

Memory

The memory you can actually use

This is why the number on the box is not the number you get. The desktop, the compositor and the graphics context never give their share back, and a model that does not leave them room does not load — it crashes.

How the memory on the box turns into memory you can use
Memory GiB
Memory on the box 24
Reserved for the system −0.80
Left for the model 23.20
Speed

Memory bandwidth

Two different limits, and people confuse them constantly. While the model is writing its answer, bandwidth rules: every single token means reading the weights again, so tokens per second is roughly bandwidth divided by model size. While the model is reading your document, compute rules: the whole input is processed at once. That is why a laptop with no GPU can take minutes before the first word appears and then type at a tolerable pace.

1,008 GB/sVendor declared · checked against the manufacturer’s page on 2026-08-25 (Source).

Compute: 82.6 TFLOPS (FP32 shader TFLOPS) — Vendor declared, under a different name · the manufacturer publishes this number as FP32 shader TFLOPS, not as FP16; checked against the manufacturer’s page on 2026-08-25 (Source). The figures in this line do not all come from the same measure across machines, because each manufacturer publishes a different one. Compare them between machines only with that in mind.

Fit

Which models fit here, and which do not

Fitting means three things at once fit in the memory left over: the model file, its context memory, and the runtime’s working headroom. Context memory is computed from each model’s declared attention pattern layer by layer, not from the generic formula — which is why several models fit here that other calculators say do not.

Every measured model against this machine: size, whether it fits at each context, and the largest context it holds
Model Compression shown Model size 8k 32k 128k Largest context that fits
qwen3-1.7b Q4_K_M 1.03 GiB yes yes 32k
qwen3-4b-2507 Q4_K_M 2.33 GiB yes yes no 64k
gemma4-e2b Q4_K_M 2.89 GiB yes yes yes 128k
gemma4-e4b Q4_K_M 4.63 GiB yes yes yes 128k
qwen3-8b Q4_K_M 4.68 GiB yes yes 32k
gemma4-12b Q4_K_M 6.63 GiB yes yes yes 128k
gpt-oss-20b Q4_K_M 10.83 GiB yes yes yes 128k
mistral-small-24b Q4_K_M 13.35 GiB yes yes no 32k
qwen3.8-27b Q4_K_M 15.33 GiB yes yes no 64k
qwen3.6-27b Q4_K_M 15.66 GiB yes yes no 64k
gemma4-26b-a4b Q4_K_M 15.78 GiB yes yes no 64k
gemma4-31b Q4_K_M 17.07 GiB yes no no 16k
qwen3-coder-30b Q4_K_M 17.28 GiB yes yes no 32k
qwen3.6-35b-a3b Q4_K_M 20.61 GiB yes yes no 32k
gpt-oss-120b Q4_K_M 58.46 GiB no no no
llama4-scout-17b Q4_K_M 60.87 GiB no no no
deepseek-v4-flash IQ4_XS 127.28 GiB no no no
glm-5.2 Q4_K_M 433.83 GiB no no no

Sizes are the real published files at the reference compression, one step per row: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Every row states which one it is showing. Q4_K_M is our quality floor: below it the loss is audible in the answers. A model that would only fit here at a harsher compression is listed as not fitting, on purpose. A dash in a context column means the model itself does not offer that context, so there is nothing to fit.

Models that fit — 14 of the 18 models measured

Models that do not fit — 4 of the 18 models measured