Home›Hardware›Xiaomi AI Cube, 80 GB — engineering prototype
Xiaomi AI Cube, 80 GB — engineering prototype
mini-PC with NPU80 GiB341 GB/s40.0 TFLOPS
You cannot buy this machine. Everything else in this hub is a product on sale; this one is an engineering prototype with no price and no release date. It is measured here because the figures going around about it are wrong, and the only way to correct them is to run it through the same arithmetic as the rest.
This machine leaves 77.00 GiB for the model once the system has its share, and it takes 16 of the 18 models we measure at the reference compression. What decides how fast it answers is not the chip: it is the 341 GB/s of memory bandwidth.
Left for the model
Estimated
Memory bandwidth
Estimated
Models that fit
Measured by Local AI Scope
The memory you can actually use
Unified memory is shared with the operating system and everything you have open. The reserve here is deliberately conservative; you can change it in the tool.
| Memory | GiB |
|---|---|
| Memory on the box | 80 |
| Reserved for the system | −3.00 |
| Left for the model | 77.00 |
The 160 GB being repeated everywhere is not what this machine has. It is the maximum the D100 chip supports. The unit Xiaomi demonstrated carries 80 GB, and 80 is what every figure on this page is computed from. The difference is not cosmetic: with 160 GB this machine would take 17 of the 18 models instead of 16, and the largest one it holds would be deepseek-v4-flash at 127.28 GiB instead of llama4-scout-17b at 60.87 GiB.
Memory bandwidth
Two different limits, and people confuse them constantly. While the model is writing its answer, bandwidth rules: every single token means reading the weights again, so tokens per second is roughly bandwidth divided by model size. While the model is reading your document, compute rules: the whole input is processed at once. That is why a laptop with no GPU can take minutes before the first word appears and then type at a tolerable pace.
341 GB/s — Estimated · the manufacturer does not publish this figure at all; ours is a conservative estimate, not a specification.
The 1.22 TB/s that went around the world is not this machine’s memory bandwidth. It is the near-memory bandwidth Xiaomi quotes for the O100 accelerator, over DRAM stacked on the compute die — a small pool, not the 80 GB where a 120B model actually sits. The bandwidth of those 80 GB is the number nobody has published, so the figure on this page is our own assumption and is labelled as one. The 200 TOPS also being repeated belongs to the O3’s NPU and is 8-bit integer throughput: another unit, on another chip. We do not convert it into the compute line above.
Compute: 40.0 TFLOPS (estimate for this class of machine) — Estimated · the manufacturer does not publish this figure at all; ours is a conservative estimate, not a specification. The figures in this line do not all come from the same measure across machines, because each manufacturer publishes a different one. Compare them between machines only with that in mind.
Which models fit here, and which do not
Fitting means three things at once fit in the memory left over: the model file, its context memory, and the runtime’s working headroom. Context memory is computed from each model’s declared attention pattern layer by layer, not from the generic formula — which is why several models fit here that other calculators say do not.
| Model | Compression shown | Model size | 8k | 32k | 128k | Largest context that fits |
|---|---|---|---|---|---|---|
| qwen3-1.7b | Q4_K_M |
1.03 GiB | yes | yes | — | 32k |
| qwen3-4b-2507 | Q4_K_M |
2.33 GiB | yes | yes | yes | 256k |
| gemma4-e2b | Q4_K_M |
2.89 GiB | yes | yes | yes | 128k |
| gemma4-e4b | Q4_K_M |
4.63 GiB | yes | yes | yes | 128k |
| qwen3-8b | Q4_K_M |
4.68 GiB | yes | yes | — | 32k |
| gemma4-12b | Q4_K_M |
6.63 GiB | yes | yes | yes | 256k |
| gpt-oss-20b | Q4_K_M |
10.83 GiB | yes | yes | yes | 128k |
| mistral-small-24b | Q4_K_M |
13.35 GiB | yes | yes | yes | 128k |
| qwen3.8-27b | Q4_K_M |
15.33 GiB | yes | yes | yes | 256k |
| qwen3.6-27b | Q4_K_M |
15.66 GiB | yes | yes | yes | 256k |
| gemma4-26b-a4b | Q4_K_M |
15.78 GiB | yes | yes | yes | 256k |
| gemma4-31b | Q4_K_M |
17.07 GiB | yes | yes | yes | 256k |
| qwen3-coder-30b | Q4_K_M |
17.28 GiB | yes | yes | yes | 256k |
| qwen3.6-35b-a3b | Q4_K_M |
20.61 GiB | yes | yes | yes | 256k |
| gpt-oss-120b | Q4_K_M |
58.46 GiB | yes | yes | yes | 128k |
| llama4-scout-17b | Q4_K_M |
60.87 GiB | yes | yes | no | 64k |
| deepseek-v4-flash | IQ4_XS |
127.28 GiB | no | no | no | — |
| glm-5.2 | Q4_K_M |
433.83 GiB | no | no | no | — |
Sizes are the real published files at the reference compression, one step per row: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Every row states which one it is showing. Q4_K_M is our quality floor: below it the loss is audible in the answers. A model that would only fit here at a harsher compression is listed as not fitting, on purpose. A dash in a context column means the model itself does not offer that context, so there is nothing to fit.
Models that fit — 16 of the 18 models measured
- qwen3-1.7b
Q4_K_M— context 32k tokens - qwen3-4b-2507
Q4_K_M— context 256k tokens - gemma4-e2b
Q4_K_M— context 128k tokens - gemma4-e4b
Q4_K_M— context 128k tokens - qwen3-8b
Q4_K_M— context 32k tokens - gemma4-12b
Q4_K_M— context 256k tokens - gpt-oss-20b
Q4_K_M— context 128k tokens - mistral-small-24b
Q4_K_M— context 128k tokens - qwen3.8-27b
Q4_K_M— context 256k tokens - qwen3.6-27b
Q4_K_M— context 256k tokens - gemma4-26b-a4b
Q4_K_M— context 256k tokens - gemma4-31b
Q4_K_M— context 256k tokens - qwen3-coder-30b
Q4_K_M— context 256k tokens - qwen3.6-35b-a3b
Q4_K_M— context 256k tokens - gpt-oss-120b
Q4_K_M— context 128k tokens - llama4-scout-17b
Q4_K_M— context 64k tokens
Models that do not fit — 2 of the 18 models measured
- deepseek-v4-flash
IQ4_XS— short by 51.29 GiB - glm-5.2
Q4_K_M— short by 372.13 GiB