Home›Guides›Mac mini M6 vs Xiaomi AI Cube: what actually fits, and in which language it runs out first

Choose · Guide

Mac mini M6 vs Xiaomi AI Cube: what actually fits, and in which language it runs out first

gpt-oss-120b is a 58.46 GiB file. The new Mac mini stops at 32 GB and misses it by 30.28 GiB; the AI Cube holds it but is a prototype you cannot buy. Measured against 20 models, and against the token cost of four languages.

Measured and written by Local AI ScopePublished Figures measured by Local AI Scope

Two machines pitched as boxes for running large language models at home, announced within forty-eight hours of each other: for local-AI nerds like us, that was a very good week. Apple’s Mac mini with M6 arrived on 25 August 2026, and Xiaomi’s AI Cube the day before. Both were talked about in terms of a 120-billion-parameter model. So we did with them exactly what we do with every other machine here: the same arithmetic, no special treatment.

Mac mini with M6 and 32 GBMac mini with M6 and 32 GBHardware sheet

gpt-oss-120b at Q4_K_M is a 58.46 GiB file. A fully specified Mac mini with M6 leaves 29.00 GiB for the model once the system has taken its share. It is 30.28 GiB short even at the smallest context step, and no context setting closes that gap, because the weights alone do not fit. Think bus and single-car garage: how carefully you reverse in has nothing to do with it. The AI Cube leaves 77.00 GiB and holds the model comfortably — the catch being that it is an engineering prototype you cannot buy.

Mac with M4 Max and 64 GBMac with M4 Max and 64 GBHardware sheet Mac mini with M5 Pro and 64 GBMac mini with M5 Pro and 64 GBHardware sheet

That is the whole comparison, honestly. Everything below is about why the numbers doing the rounds say something else.

The Mac mini number that matters is 32, not 899 and not 2

The figures that made the headlines were the price ($899) and the process (Apple’s first 2-nanometer chip). Nice to know, but neither one decides whether a model runs. What decides it is the memory ceiling, and Apple’s own announcement spells it out: the M6 ships with “16GB of standard unified memory configurable up to 32GB”.

Mac mini with M6 connected to a display
Mac mini with M6 connected to a display Photo: Apple

So let us say it clearly: there is no Mac mini with M6 and 64 GB. Apple announced a second Mac mini the same day — the one with M5 Pro, from $1,699 — and that is the machine that goes up to 64 GB, at 307 GB/s. Different chip, different ceiling, roughly twice the price. If you saw “new Mac mini” and “64 GB” in the same sentence, the sentence was mixing up two products.

Here is what that ceiling costs you, measured against the 20 models in our catalogue at the reference compression:

Machine Memory left for the model Models that fit at 32k Largest model it holds Ceiling for gpt-oss-120b
Mac mini, M6, 32 GB 29.00 GiB 15 of 20 qwen3.6-35b-a3b, 20.61 GiB does not fit
Mac, M4 Max, 64 GB 61.00 GiB 16 of 20 gpt-oss-120b, 58.46 GiB 32k tokens
Xiaomi AI Cube, 80 GB 77.00 GiB 17 of 20 llama4-scout-17b, 60.87 GiB 128k tokens

Two things here deserve a second read. First, going from 32 GB to 64 GB buys exactly one extra model out of 20 — but that one model is the 120B everybody was talking about, which is why the ceiling matters more than the count. Second, at 64 GB the 120B does fit, but only just: at 32k of context it is using 60.79 of the 61.00 GiB available, and the next step up, 64k, would need 1.51 GiB more than the machine has. On the AI Cube the same model reaches 128k — its own declared maximum — with 11.04 GiB still unused, and that is the point where the limit stops being the machine and starts being the model. Memory does not run out at the file; it runs out at the conversation.

The Xiaomi figures being repeated belong to other things

Two numbers travelled a lot further than the machine itself, and neither of them describes it.

Xiaomi AI Cube, front view
Xiaomi AI Cube, front view Photo: Xiaomi

1.22 TB/s is not the AI Cube’s memory bandwidth. It is the near-memory bandwidth Xiaomi quotes for the Xring O100 accelerator, over DRAM stacked directly on the compute die. Picture a paddling pool right next to the maths unit, not the 80 GB lake where a 120B model actually sits. The bandwidth of those 80 GB is the number nobody has published — so the figure on our sheet for it is our own assumption, and it is labelled as an estimate rather than dressed up as a specification.

Opens on YouTube (no cookies) only when you press play. More videos on its sheet →

160 GB is not what the machine has. It is the maximum the Xring D100 chip supports. The unit Xiaomi demonstrated carries 80 GB, and 80 is what every figure above is worked out from. And no, the difference is not cosmetic: at 160 GB this machine would take 19 of the 20 models instead of 17, and the largest one it held would be deepseek-v4-flash at 127.28 GiB rather than llama4-scout-17b at 60.87 GiB.

A third number, 200 TOPS, belongs to the NPU inside the Xring O3 and is 8-bit integer throughput: another unit, on another one of the three chips. We do not convert it into anything.

We measured it anyway, because wrong figures about it are already out there and the only way to correct them is to do the arithmetic. But its sheet carries the warning in its title, not in a footnote, because everything else in that section is a machine you can order today.

And in German, the ceiling arrives sooner

This is the part we love digging into. A context ceiling is measured in tokens, and the same document is not the same number of tokens in every language. On the tokenizer gpt-oss-120b actually uses, we measured a parallel corpus of the same texts in four languages. German costs 28.6% more tokens than English for the same content overall, and 41.9% more on code.

Kind of text English Spanish French German
Code ×1.00 ×1.20 ×1.32 ×1.42
Support ticket ×1.00 ×1.20 ×1.36 ×1.26
Contract ×1.00 ×1.14 ×1.23 ×1.25
Email ×1.00 ×1.12 ×1.19 ×1.17

To be straight about what this is: the cost in tokens of saying the same thing, and nothing else. It says nothing about how well any model answers in any language; we have no measurement of that, so we do not claim one. It is also why you will never see us publish “tokens per word” as a comparison: German says the same thing in fewer, longer words, so that ratio flatters it. The fair comparison is the token count of identical content, which is exactly what the table above is.

What we could not verify

The evidence label on every figure here sits on its machine sheet, but the gaps deserve to be said out loud:

  • The AI Cube’s memory bandwidth and its compute figure are ours, not Xiaomi’s. Xiaomi has published neither for the 80 GB pool. Its own announcement lives on a social post that requires an account to read, so there is no manufacturer page to cite. Both figures are marked as estimates on the sheet.
  • Apple publishes no GPU TFLOPS for any M chip. The compute figure on the Mac mini sheet is our own estimate, on the same per-GPU-core basis as the other Apple machines here. That column is not comparable between manufacturers at all, and every sheet says so.
  • Apple calls the M6 its “first state-of-the-art 2-nanometer chip” without naming a foundry. The specific process repeated in the coverage is not Apple’s word, so it is not ours either.
  • The 80 GB itself comes from the reporting of Xiaomi’s event, not from a Xiaomi specification page. It is consistent across the outlets that were there; it is not a document we can link.

The memory a model needs, the cache it needs at each context and the token cost of each language are measured here, on the files and the corpus. The two figures above that we could not confirm are marked as estimates everywhere they appear, and neither of them changes a single verdict about what fits: fit is decided by memory capacity, and capacity is the one number both manufacturers do publish.

So which one

Apple M4, M4 Pro and M4 Max chips
Apple M4, M4 Pro and M4 Max chips Image: Apple

If a 120B is not the requirement, the Mac mini with M6 takes 15 of the 20 models we measure, up to qwen3.6-35b-a3b at 20.61 GiB, and that covers most of what people actually run locally. Plenty to have fun with.

Which of the 20 fits on the machine you already own, at the context your work needs and in the language you work in, is what the Model Finder answers. The full comparison of all 20 models and all 18 machines is a table each, and every figure in them carries its evidence label.