Hardware
Every hardware sheet answers the question the spec sheet dodges: how much memory is really left once the system takes its share, and which of the 18 measured models fit in what remains — with the context each one can actually hold. Two lists per machine, the ones that fit and the ones that do not, so you can see what not to buy.
| Machine | Memory left for the model | Bandwidth | Models that fit at 8k | and at 32k | Largest model it takes |
|---|---|---|---|---|---|
| PC with an RTX 4060 8 GB | 7.20 GiB | 272 GB/s derived, not published | 5 / 18 | 3 / 18 | qwen3-8b 4.68 GiB |
| PC with an RTX 3060 12 GB | 11.20 GiB | 360 GB/s derived, not published | 6 / 18 | 6 / 18 | gemma4-12b 6.63 GiB |
| PC with an RTX 4070 12 GB | 11.20 GiB | 504 GB/s derived, not published | 6 / 18 | 6 / 18 | gemma4-12b 6.63 GiB |
| Mac with M4 and 16 GB unified memory | 13.00 GiB | 120 GB/s vendor declared | 7 / 18 | 7 / 18 | gpt-oss-20b 10.83 GiB |
| Office laptop, CPU only, 16 GB | 13.00 GiB | 83 GB/s estimated | 7 / 18 | 7 / 18 | gpt-oss-20b 10.83 GiB |
| Mac with M4 Pro and 24 GB | 21.00 GiB | 273 GB/s vendor declared | 13 / 18 | 11 / 18 | qwen3-coder-30b 17.28 GiB |
| Workstation with an RTX 4090 24 GB | 23.20 GiB | 1008 GB/s vendor declared | 14 / 18 | 13 / 18 | qwen3.6-35b-a3b 20.61 GiB |
| Mini-PC with an NPU and 32 GB LPDDR5X | 29.00 GiB | 120 GB/s estimated | 14 / 18 | 14 / 18 | qwen3.6-35b-a3b 20.61 GiB |
| Server with 2× RTX 3090 (48 GB) | 47.20 GiB | 936 GB/s vendor declared | 14 / 18 | 14 / 18 | qwen3.6-35b-a3b 20.61 GiB |
| Mac with M4 Max and 64 GB | 61.00 GiB | 546 GB/s vendor declared | 15 / 18 | 15 / 18 | gpt-oss-120b 58.46 GiB |
Fitting means the model file, its context memory and the runtime headroom all fit in what is left after the system takes its share, at Q4_K_M or the nearest compression the author publishes. Bandwidth carries its provenance in the cell: five of these ten figures are not published by the manufacturer anywhere, and we say so rather than passing them off as specifications.