Home›Guides›What AI your computer can run, and what you give up to run it

Choose · Guide

What AI your computer can run, and what you give up to run it

On an 8 GB card, 6 of the 20 models we measure fit alongside an eight-page document, and the biggest is 4.98 GiB. The largest model we catalogue is 433.83 GiB. That is 87x, and it is the honest starting point.

Measured and written by Local AI ScopePublished Figures measured by Local AI Scope

Your computer can almost certainly run a local AI. What it cannot run is the one you have been using. On an 8 GB graphics card, 6 of the 20 models we measure fit alongside an eight-page document, and the biggest of them is 4.98 GiB; the largest model we catalogue is 433.83 GiB. Everything else on this page follows from that gap.

That is not a reason to skip local AI. It is the reason to know what you are trading before you spend a weekend installing something.

What actually changes when the model runs on your machine

Four things change, and only four.

Nothing you type leaves the machine. Not the document, not the prompt, not the draft you did not send. There is no processor agreement to sign, no international-transfer clause to read, no training opt-out to trust. For anything covered by GDPR, that removes the paperwork rather than mitigating it.

The cost stops being per use. You pay once for hardware and electricity. There is no meter running while you paste a long contract in and ask for a summary five times.

It works with the network unplugged. Once the file is on disk, a flight, a rural office or a locked-down network stops being a blocker.

And it is smaller and slower than what you are used to. That is the part nobody puts next to the price, so the rest of this page puts numbers on it.

How we count, before any number

Every figure here uses one compression per model: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Q4_K_M is our quality floor, because below it the loss shows in the answers.

That choice has a consequence worth stating plainly: a model that would only fit on your machine at a harsher compression is counted here as not fitting. We could publish a bigger number by letting the compression slide. We do not, and it is the same rule the hardware pages and the model pages use, so the counts on this site agree with each other.

How much smaller, exactly

A local model is a file, and it has to fit in the memory left after the operating system takes its share — together with the memory the conversation itself occupies and the runtime’s own working margin. We reserve 0.8 GiB on dedicated graphics cards and 3 GiB on unified or system memory. On PCs with Ryzen AI Max+ a different limit applies: AMD lets the GPU use at most 96 GiB of 128 GB and 160 GiB of 192 GB.

For a 4,000-word document in English:

Machine Memory for the model Models that fit Biggest that fits Size Distance to glm-5.2
PC with an RTX 4060 8 GB 7.20 GiB 6 of 20 granite-4.2-8b 4.98 GiB 87x
PC with an RTX 3060 12 GB 11.20 GiB 7 of 20 gemma4-12b 6.63 GiB 65x
PC with an RTX 4070 12 GB 11.20 GiB 7 of 20 gemma4-12b 6.63 GiB 65x
Mac with M4 and 16 GB unified 13.00 GiB 8 of 20 gpt-oss-20b 10.83 GiB 40x
Laptop, 16 GB, no graphics card 13.00 GiB 8 of 20 gpt-oss-20b 10.83 GiB 40x
Mac mini with M6 and 16 GB 13.00 GiB 8 of 20 gpt-oss-20b 10.83 GiB 40x
Mac with M4 Pro and 24 GB 21.00 GiB 14 of 20 qwen3-coder-30b 17.28 GiB 25x
Workstation with an RTX 4090 24 GB 23.20 GiB 15 of 20 qwen3.6-35b-a3b 20.61 GiB 21x
Mini-PC with an NPU and 32 GB 29.00 GiB 15 of 20 qwen3.6-35b-a3b 20.61 GiB 21x
Mac mini with M6 and 32 GB 29.00 GiB 15 of 20 qwen3.6-35b-a3b 20.61 GiB 21x
Server with 2× RTX 3090 (48 GB) 47.20 GiB 15 of 20 qwen3.6-35b-a3b 20.61 GiB 21x
Mac with M4 Max and 64 GB 61.00 GiB 16 of 20 gpt-oss-120b 58.46 GiB 7x
Mac mini with M5 Pro and 64 GB 61.00 GiB 16 of 20 gpt-oss-120b 58.46 GiB 7x
Xiaomi AI Cube, 80 GB (prototype) 77.00 GiB 17 of 20 llama4-scout-17b 60.87 GiB 7x
PC with Ryzen AI Max+ 395 and 128 GB 96.00 GiB 18 of 20 qwen3.8-flash-next 87.25 GiB 5x
Mac Studio with M5 Max and 128 GB 125.00 GiB 18 of 20 qwen3.8-flash-next 87.25 GiB 5x
PC with Ryzen AI Max+ PRO 495 and 192 GB 160.00 GiB 19 of 20 deepseek-v4-flash 127.28 GiB 3x
Mac Studio with M5 Ultra and 512 GB 509.00 GiB 20 of 20 glm-5.2 433.83 GiB 1x

Calculated by Local AI Scope from published file sizes and each model’s real attention pattern, every row at the same reference compression so the sizes are comparable. Sizes come from the authors’ own published builds: glm-5.2, gemma4-12b, gpt-oss-20b. glm-5.2 is 433.83 GiB in Q4_K_M; the least compressed build we catalogue (Q8_0) is 746.32 GiB.

Two things in that table matter more than the ratio. The first is that the ladder reaches parity only at the very top: a 64 GB Mac, which is not a cheap machine, is still 7 times short of the largest weights we catalogue, 128 GB only brings that down to 5, 192 GB to 3, and the only machine here that reaches 20 of 20 is the Mac Studio with M5 Ultra and 512 GB. The second is that the step from 8 GB to 24 GB buys far more than the step from 24 GB to 64 GB — 6 models to 15, then 15 to 16.

How slow is slow

Speed splits into two numbers that behave differently, and conflating them is why people are surprised.

Writing the answer is limited by memory bandwidth. Reading your document first — before a single word comes back — is limited by raw compute. Short chat hides the second number entirely; a long document is nothing but the second number.

Same 4,000-word English document, best fitting model for that job on each machine:

Machine Model Writes at Waits before the first word
Laptop, 16 GB, no graphics card gemma4-12b ~4.1 tokens/s ~7 min 48 s
Mac with M4 and 16 GB unified gemma4-12b ~9.3 tokens/s ~87 s
PC with an RTX 3060 12 GB gemma4-12b ~33 tokens/s ~25 s
PC with an RTX 4060 8 GB qwen3-8b ~35 tokens/s ~15 s
Workstation with an RTX 4090 24 GB qwen3.6-35b-a3b ~213 tokens/s ~2 s

Estimated by Local AI Scope, not measured: the formula uses each machine’s declared memory bandwidth and compute, with conservative efficiency assumptions. It is published so you can check it, and it never decides whether a model fits.

Now the same machines with a normal chat message of about 1,000 words. The RTX 4060 waits around 4 seconds and writes at roughly 33 tokens/s. The 16 GB laptop with no graphics card waits about 16 seconds and writes at roughly 17 tokens/s.

So the honest version is not “local AI is slow”. It is: chat is fine on modest hardware, and long documents are where it falls apart. If your work is email, notes and questions, the machine you already own is probably enough. If your work is contracts, case files or long reports, the wait before the first word is what will make you give up, and it is a compute problem that more memory alone does not fix.

Is a local AI good enough to replace ChatGPT for my day-to-day work?

Split the question, because the two halves have different answers.

Short jobs that fit in a page or two — drafting, rewriting, summarising, questions about a document, working offline: these are the shape of job a model on a normal machine is built for, and the privacy and the flat cost are real gains.

Long-document reasoning, code across a whole repository, anything where you were leaning on a frontier model: short of a 512 GB Mac Studio, the file on your disk is 3 to 87 times smaller than the largest one we catalogue, depending on your machine, and no setting closes that gap.

What that size difference costs you in answer quality, we have not measured, and we are not going to put a number on it. We have not measured answer quality by language either. Measuring it means running the models and grading what they reply, which needs a GPU we do not have. We measure what fits, what a language costs in tokens, and what the privacy decision implies. Anyone who tells you their local model is “as good as GPT in German” without publishing how they measured it is guessing.

One thing that does change the answer: the language you work in

The same text is not the same number of tokens in every language, and context is reserved in tokens. Measured with each model’s own tokenizer over a parallel corpus of ours (sha256[:16] 2f6c96b6ea0b161f), a contract costs up to 1.76x more tokens in German than in English, up to 1.47x in French and up to 1.39x in Spanish, depending on the model.

That is not a rounding error, because the runtime reserves context in powers of two: cross a step and memory jumps. On an RTX 4060, a German document starts costing you a model at 2,800 words, a Spanish one at 4,200 and a French one at 4,200. On the 16 GB machines the same crossing happens at 11,300 words in German, 16,800 in Spanish and 16,600 in French. The token calculator has the per-model figures.

If you are going to buy something, buy memory

Two rules survive every scenario we ran.

  • Memory decides what fits; bandwidth decides how fast it runs. An RTX 4070 12 GB and an RTX 3060 12 GB fit exactly the same models. The 4070 runs them faster and adds none.
  • The file size is the specification, not the parameter count. gemma4-12b is 6.63 GiB and qwen3-coder-30b is 17.28 GiB: the second will not fit on any 12 GB card no matter what its name suggests.

The hardware pages carry the two lists for each machine — what fits and what does not — with the memory and bandwidth each one actually has, counted the same way as this page. We calculate that list before we look at what is on sale, and we say so in our editorial policy: a machine never ranks higher here because it is easier to buy.

Three questions, and the tool answers them

The model finder needs three answers and runs entirely in your browser — what you type into it never reaches us.

  1. What machine do you have? Memory and type, not brand. Start from your own: RTX 4060 8 GB · RTX 3060 12 GB · Mac M4 16 GB · laptop with no graphics card · RTX 4090 24 GB.
  2. What are you feeding it, how long, and in what language? A 600-word email and a 12,000-word contract are different machines’ problems, and the language moves the answer.
  3. What is in that data? Nothing sensitive, company-internal, personal data, or must-not-leave. That last one is the only answer that makes the verdict a requirement rather than a preference.

One difference worth knowing: the tool will also show you what fits if you allow harsher compression than Q4_K_M, and it labels those clearly. The counts on this page never do.