Measured, not estimated

Which local AI can you actually run?

Tell it what machine you have, what you work with and how sensitive that work is. It answers with measurements: what fits, how much context you can afford, how long you will wait, and whether keeping it local is the right call at all.

Find a setup for my machine

The catalogue Measured
Models with a sheet18
Smallest / largest1.03 / 433.83 GiB
Fit in 8 GB5 of 18
Costliest languageDE ×1.62
Measured on2026-08-25
18models measured
10machine profiles
5runtimes
4languages
Why this site exists

Three numbers that no VRAM calculator shows you

Each one is reproducible from the published method, and each one changes a decision somebody is about to make.

×5.2

How much the generic VRAM formula overestimates the memory of a 2026 model at 32k context. Most calculators still use it, so they tell people something does not fit when it does.

Measured by Local AI Scope
×1.62

What the same document costs in German versus English, in tokens, with the Qwen3 family. You pay for tokens in context, in memory and in your cloud bill.

Measured by Local AI Scope
30.6%

Of 640 tested situations where changing only the language changes the amount of context you must reserve. In 14.4% it changes the recommended model outright.

Measured by Local AI Scope
Provenance

Where every number comes from

Each model’s size is that of the real file published by its author. Context memory is calculated from the attention pattern declared layer by layer in the official configuration, not from the generic formula: in gemma-4-12B at 32k, that formula gives 12.00 GiB where the real memory is 2.31 GiB. The cost of each language is measured by running every model’s own tokenizer over a parallel corpus — the same content in all four languages.

What is not measured yet: whether a model writes better in French or German. That requires running and evaluating them, and until that measurement exists this tool does not claim it. Speeds are estimates from manufacturer specifications; they never decide whether something fits.

Measured by Local AI ScopeCommunity verifiedEstimatedVendor declaredStale

Read the full method

Questions

What people ask before they trust a number here

Why does your calculator say a model fits when others say it does not?

Because the context cache is computed from the attention pattern each model declares layer by layer, not from the textbook formula. gemma4-12b declares 48 layers and only 8 of them are full attention; the other 40 slide over 1024 tokens and never hold more. At 32k that is 2.31 GiB of real cache against the 12.00 GiB the generic formula returns.

Does the language I work in really change which model I should run?

Yes, and it is measured, not argued. The same 12,000-word contract costs ×1.62 more tokens in German than in English with the Qwen3 family. In 30.6 % of the 640 situations we tested, changing only the language changes the amount of context you have to reserve; in 14.4 % it changes the recommended model outright.

Do you know whether a model writes better in French or in German?

No, and this site does not claim it. Measuring answer quality per language means running and evaluating the models, which we have not done. What is measured here is what a language costs in tokens — and tokens are what you pay for in context, in memory and in a cloud bill. Anything about quality would be an opinion wearing a number.

Does anything I type into the tool leave my browser?

No. The Model Finder computes everything locally from a dataset the page has already downloaded. There is no account, no email field and no request carrying your answers. The two typefaces are self-hosted for the same reason: nothing on this page asks a third party for anything.

One question, one answer

Find out what your own machine can run

Ten machine profiles, eighteen models, four languages. It answers in your browser and tells you where each number came from.

Find a setup for my machine