Your computer can almost certainly run a local AI. What it cannot run is the one you have been using. On an 8 GB graphics card, 6 of the 20 models we measure fit alongside an eight-page document, and the biggest of them is 4.98 GiB; the largest model we catalogue is 433.83 GiB. Everything else on this page follows from that gap.
That is not a reason to skip local AI. It is the reason to know what you are trading before you spend a weekend installing something.
What actually changes when the model runs on your machine
Four things change, and only four.
Nothing you type leaves the machine. Not the document, not the prompt, not the draft you did not send. There is no processor agreement to sign, no international-transfer clause to read, no training opt-out to trust. For anything covered by GDPR, that removes the paperwork rather than mitigating it.
The cost stops being per use. You pay once for hardware and electricity. There is no meter running while you paste a long contract in and ask for a summary five times.
It works with the network unplugged. Once the file is on disk, a flight, a rural office or a locked-down network stops being a blocker.
And it is smaller and slower than what you are used to. That is the part nobody puts next to the price, so the rest of this page puts numbers on it.
How we count, before any number
Every figure here uses one compression per model: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Q4_K_M is our quality floor, because below it the loss shows in the answers.
That choice has a consequence worth stating plainly: a model that would only fit on your machine at a harsher compression is counted here as not fitting. We could publish a bigger number by letting the compression slide. We do not, and it is the same rule the hardware pages and the model pages use, so the counts on this site agree with each other.
How much smaller, exactly
A local model is a file, and it has to fit in the memory left after the operating system takes its share — together with the memory the conversation itself occupies and the runtime’s own working margin. We reserve 0.8 GiB on dedicated graphics cards and 3 GiB on unified or system memory. On PCs with Ryzen AI Max+ a different limit applies: AMD lets the GPU use at most 96 GiB of 128 GB and 160 GiB of 192 GB.
For a 4,000-word document in English:
| Machine | Memory for the model | Models that fit | Biggest that fits | Size | Distance to glm-5.2 |
|---|---|---|---|---|---|
| PC with an RTX 4060 8 GB | 7.20 GiB | 6 of 20 | granite-4.2-8b |
4.98 GiB | 87x |
| PC with an RTX 3060 12 GB | 11.20 GiB | 7 of 20 | gemma4-12b |
6.63 GiB | 65x |
| PC with an RTX 4070 12 GB | 11.20 GiB | 7 of 20 | gemma4-12b |
6.63 GiB | 65x |
| Mac with M4 and 16 GB unified | 13.00 GiB | 8 of 20 | gpt-oss-20b |
10.83 GiB | 40x |
| Laptop, 16 GB, no graphics card | 13.00 GiB | 8 of 20 | gpt-oss-20b |
10.83 GiB | 40x |
| Mac mini with M6 and 16 GB | 13.00 GiB | 8 of 20 | gpt-oss-20b |
10.83 GiB | 40x |
| Mac with M4 Pro and 24 GB | 21.00 GiB | 14 of 20 | qwen3-coder-30b |
17.28 GiB | 25x |
| Workstation with an RTX 4090 24 GB | 23.20 GiB | 15 of 20 | qwen3.6-35b-a3b |
20.61 GiB | 21x |
| Mini-PC with an NPU and 32 GB | 29.00 GiB | 15 of 20 | qwen3.6-35b-a3b |
20.61 GiB | 21x |
| Mac mini with M6 and 32 GB | 29.00 GiB | 15 of 20 | qwen3.6-35b-a3b |
20.61 GiB | 21x |
| Server with 2× RTX 3090 (48 GB) | 47.20 GiB | 15 of 20 | qwen3.6-35b-a3b |
20.61 GiB | 21x |
| Mac with M4 Max and 64 GB | 61.00 GiB | 16 of 20 | gpt-oss-120b |
58.46 GiB | 7x |
| Mac mini with M5 Pro and 64 GB | 61.00 GiB | 16 of 20 | gpt-oss-120b |
58.46 GiB | 7x |
| Xiaomi AI Cube, 80 GB (prototype) | 77.00 GiB | 17 of 20 | llama4-scout-17b |
60.87 GiB | 7x |
| PC with Ryzen AI Max+ 395 and 128 GB | 96.00 GiB | 18 of 20 | qwen3.8-flash-next |
87.25 GiB | 5x |
| Mac Studio with M5 Max and 128 GB | 125.00 GiB | 18 of 20 | qwen3.8-flash-next |
87.25 GiB | 5x |
| PC with Ryzen AI Max+ PRO 495 and 192 GB | 160.00 GiB | 19 of 20 | deepseek-v4-flash |
127.28 GiB | 3x |
| Mac Studio with M5 Ultra and 512 GB | 509.00 GiB | 20 of 20 | glm-5.2 |
433.83 GiB | 1x |
Calculated by Local AI Scope from published file sizes and each model’s real attention pattern, every row at the same reference compression so the sizes are comparable. Sizes come from the authors’ own published builds: glm-5.2, gemma4-12b, gpt-oss-20b. glm-5.2 is 433.83 GiB in Q4_K_M; the least compressed build we catalogue (Q8_0) is 746.32 GiB.
Two things in that table matter more than the ratio. The first is that the ladder reaches parity only at the very top: a 64 GB Mac, which is not a cheap machine, is still 7 times short of the largest weights we catalogue, 128 GB only brings that down to 5, 192 GB to 3, and the only machine here that reaches 20 of 20 is the Mac Studio with M5 Ultra and 512 GB. The second is that the step from 8 GB to 24 GB buys far more than the step from 24 GB to 64 GB — 6 models to 15, then 15 to 16.
How slow is slow
Speed splits into two numbers that behave differently, and conflating them is why people are surprised.
Writing the answer is limited by memory bandwidth. Reading your document first — before a single word comes back — is limited by raw compute. Short chat hides the second number entirely; a long document is nothing but the second number.
Same 4,000-word English document, best fitting model for that job on each machine:
| Machine | Model | Writes at | Waits before the first word |
|---|---|---|---|
| Laptop, 16 GB, no graphics card | gemma4-12b |
~4.1 tokens/s | ~7 min 48 s |
| Mac with M4 and 16 GB unified | gemma4-12b |
~9.3 tokens/s | ~87 s |
| PC with an RTX 3060 12 GB | gemma4-12b |
~33 tokens/s | ~25 s |
| PC with an RTX 4060 8 GB | qwen3-8b |
~35 tokens/s | ~15 s |
| Workstation with an RTX 4090 24 GB | qwen3.6-35b-a3b |
~213 tokens/s | ~2 s |
Estimated by Local AI Scope, not measured: the formula uses each machine’s declared memory bandwidth and compute, with conservative efficiency assumptions. It is published so you can check it, and it never decides whether a model fits.
Now the same machines with a normal chat message of about 1,000 words. The RTX 4060 waits around 4 seconds and writes at roughly 33 tokens/s. The 16 GB laptop with no graphics card waits about 16 seconds and writes at roughly 17 tokens/s.
So the honest version is not “local AI is slow”. It is: chat is fine on modest hardware, and long documents are where it falls apart. If your work is email, notes and questions, the machine you already own is probably enough. If your work is contracts, case files or long reports, the wait before the first word is what will make you give up, and it is a compute problem that more memory alone does not fix.
Is a local AI good enough to replace ChatGPT for my day-to-day work?
Split the question, because the two halves have different answers.
Short jobs that fit in a page or two — drafting, rewriting, summarising, questions about a document, working offline: these are the shape of job a model on a normal machine is built for, and the privacy and the flat cost are real gains.
Long-document reasoning, code across a whole repository, anything where you were leaning on a frontier model: short of a 512 GB Mac Studio, the file on your disk is 3 to 87 times smaller than the largest one we catalogue, depending on your machine, and no setting closes that gap.
What that size difference costs you in answer quality, we have not measured, and we are not going to put a number on it. We have not measured answer quality by language either. Measuring it means running the models and grading what they reply, which needs a GPU we do not have. We measure what fits, what a language costs in tokens, and what the privacy decision implies. Anyone who tells you their local model is “as good as GPT in German” without publishing how they measured it is guessing.
One thing that does change the answer: the language you work in
The same text is not the same number of tokens in every language, and context is reserved in tokens. Measured with each model’s own tokenizer over a parallel corpus of ours (sha256[:16] 2f6c96b6ea0b161f), a contract costs up to 1.76x more tokens in German than in English, up to 1.47x in French and up to 1.39x in Spanish, depending on the model.
That is not a rounding error, because the runtime reserves context in powers of two: cross a step and memory jumps. On an RTX 4060, a German document starts costing you a model at 2,800 words, a Spanish one at 4,200 and a French one at 4,200. On the 16 GB machines the same crossing happens at 11,300 words in German, 16,800 in Spanish and 16,600 in French. The token calculator has the per-model figures.
If you are going to buy something, buy memory
Two rules survive every scenario we ran.
- Memory decides what fits; bandwidth decides how fast it runs. An RTX 4070 12 GB and an RTX 3060 12 GB fit exactly the same models. The 4070 runs them faster and adds none.
- The file size is the specification, not the parameter count.
gemma4-12bis 6.63 GiB andqwen3-coder-30bis 17.28 GiB: the second will not fit on any 12 GB card no matter what its name suggests.
The hardware pages carry the two lists for each machine — what fits and what does not — with the memory and bandwidth each one actually has, counted the same way as this page. We calculate that list before we look at what is on sale, and we say so in our editorial policy: a machine never ranks higher here because it is easier to buy.
Three questions, and the tool answers them
The model finder needs three answers and runs entirely in your browser — what you type into it never reaches us.
- What machine do you have? Memory and type, not brand. Start from your own: RTX 4060 8 GB · RTX 3060 12 GB · Mac M4 16 GB · laptop with no graphics card · RTX 4090 24 GB.
- What are you feeding it, how long, and in what language? A 600-word email and a 12,000-word contract are different machines’ problems, and the language moves the answer.
- What is in that data? Nothing sensitive, company-internal, personal data, or must-not-leave. That last one is the only answer that makes the verdict a requirement rather than a preference.
One difference worth knowing: the tool will also show you what fits if you allow harsher compression than Q4_K_M, and it labels those clearly. The counts on this page never do.