Home›Hardware›Mac Studio with M5 Ultra and 512 GB
Mac Studio with M5 Ultra and 512 GB
Apple silicon512 GiB1,200 GB/s36.8 TFLOPS
You cannot have this machine yet. It is announced, not on a shelf. Everything below is arithmetic on the manufacturer’s declared figures, not a machine anyone has run a model on.
Apple opened sales for the rest of the range on 22 September 2026. When we read its store, on 23 September 2026, the memory selector still gave late October for the 512 GB option, and that configuration has no page of its own: it is announced, not even orderable.
This machine leaves 509.00 GiB for the model once the system has its share, and it takes 20 of the 20 models we measure at the reference compression. What decides how fast it answers is not the chip: it is the 1,200 GB/s of memory bandwidth.
Left for the model
Estimated
Memory bandwidth
Vendor declared
Models that fit
Measured by Local AI Scope
Videos: Mac Studio with M5 Ultra and 512 GB 4
Official footage and the best walk-throughs we found. Press play and it opens right here.
The videos live on YouTube, and nothing loads from there until you press play (no-cookie mode). The thumbnail is served by us.
The memory you can actually use
Unified memory is shared with the operating system and everything you have open. The reserve here is deliberately conservative; you can change it in the tool.
| Memory | GiB |
|---|---|
| Memory on the box | 512 |
| Reserved for the system | −3.00 |
| Left for the model | 509.00 |
512 GB is the ceiling, not what you get for the starting price. It is the largest configuration Apple offers, it is a configure-to-order option whose price Apple does not publish in the announcement, and it is the one that ships last, in late October. Everything on this page is computed from 512 — so if you buy a smaller configuration, the fit table below is not yours. This is also the first machine in this catalogue where every model we measure fits, including glm-5.2 at 433.83 GiB, which no other machine here takes.
Memory bandwidth
Two different limits, and people confuse them constantly. While the model is writing its answer, bandwidth rules: every single token means reading the weights again, so tokens per second is roughly bandwidth divided by model size. While the model is reading your document, compute rules: the whole input is processed at once. That is why a laptop with no GPU can take minutes before the first word appears and then type at a tolerable pace.
1,200 GB/s — Vendor declared · checked against the manufacturer’s page on 2026-08-31 (Source).
The 1.2 TB/s belongs to the M5 Ultra. Apple announced two Mac Studios the same day: the other one carries M5 Max, a different chip that tops out at 614 GB/s and 128 GB, from $2,499. Reading the Ultra’s bandwidth onto the Max would be wrong by a factor of two. And Apple publishes no TFLOPS for any M chip — only core counts and comparisons against its own previous generation — so the compute line below is our own floor, not a specification.
Compute: 36.8 TFLOPS (estimate for this class of machine) — Estimated · the manufacturer does not publish this figure at all; ours is a conservative estimate, not a specification. The figures in this line do not all come from the same measure across machines, because each manufacturer publishes a different one. Compare them between machines only with that in mind.
Which models fit here, and which do not
Fitting means three things at once fit in the memory left over: the model file, its context memory, and the runtime’s working headroom. Context memory is computed from each model’s declared attention pattern layer by layer, not from the generic formula — which is why several models fit here that other calculators say do not.
| Model | Compression shown | Model size | 8k | 32k | 128k | Largest context that fits |
|---|---|---|---|---|---|---|
| qwen3-1.7b | Q4_K_M |
1.03 GiB | yes | yes | — | 32k |
| qwen3-4b-2507 | Q4_K_M |
2.33 GiB | yes | yes | yes | 256k |
| gemma4-e2b | Q4_K_M |
2.89 GiB | yes | yes | yes | 128k |
| gemma4-e4b | Q4_K_M |
4.63 GiB | yes | yes | yes | 128k |
| qwen3-8b | Q4_K_M |
4.68 GiB | yes | yes | — | 32k |
| granite-4.2-8b | Q4_K_M |
4.98 GiB | yes | yes | yes | 128k |
| gemma4-12b | Q4_K_M |
6.63 GiB | yes | yes | yes | 256k |
| gpt-oss-20b | Q4_K_M |
10.83 GiB | yes | yes | yes | 128k |
| mistral-small-24b | Q4_K_M |
13.35 GiB | yes | yes | yes | 128k |
| qwen3.8-27b | Q4_K_M |
15.33 GiB | yes | yes | yes | 256k |
| qwen3.6-27b | Q4_K_M |
15.66 GiB | yes | yes | yes | 256k |
| gemma4-26b-a4b | Q4_K_M |
15.78 GiB | yes | yes | yes | 256k |
| gemma4-31b | Q4_K_M |
17.07 GiB | yes | yes | yes | 256k |
| qwen3-coder-30b | Q4_K_M |
17.28 GiB | yes | yes | yes | 256k |
| qwen3.6-35b-a3b | Q4_K_M |
20.61 GiB | yes | yes | yes | 256k |
| gpt-oss-120b | Q4_K_M |
58.46 GiB | yes | yes | yes | 128k |
| llama4-scout-17b | Q4_K_M |
60.87 GiB | yes | yes | yes | 1024k |
| qwen3.8-flash-next | IQ4_XS |
87.25 GiB | yes | yes | yes | 256k |
| deepseek-v4-flash | IQ4_XS |
127.28 GiB | yes | yes | yes | 1024k |
| glm-5.2 | Q4_K_M |
433.83 GiB | yes | no | no | 16k |
Sizes are the real published files at the reference compression, one step per row: Q4_K_M where the author publishes it, the nearest neighbour where they do not. Every row states which one it is showing. Q4_K_M is our quality floor: below it the loss is audible in the answers. A model that would only fit here at a harsher compression is listed as not fitting, on purpose. A dash in a context column means the model itself does not offer that context, so there is nothing to fit.
Models that fit — 20 of the 20 models measured
- qwen3-1.7b
Q4_K_M— context 32k tokens - qwen3-4b-2507
Q4_K_M— context 256k tokens - gemma4-e2b
Q4_K_M— context 128k tokens - gemma4-e4b
Q4_K_M— context 128k tokens - qwen3-8b
Q4_K_M— context 32k tokens - granite-4.2-8b
Q4_K_M— context 128k tokens - gemma4-12b
Q4_K_M— context 256k tokens - gpt-oss-20b
Q4_K_M— context 128k tokens - mistral-small-24b
Q4_K_M— context 128k tokens - qwen3.8-27b
Q4_K_M— context 256k tokens - qwen3.6-27b
Q4_K_M— context 256k tokens - gemma4-26b-a4b
Q4_K_M— context 256k tokens - gemma4-31b
Q4_K_M— context 256k tokens - qwen3-coder-30b
Q4_K_M— context 256k tokens - qwen3.6-35b-a3b
Q4_K_M— context 256k tokens - gpt-oss-120b
Q4_K_M— context 128k tokens - llama4-scout-17b
Q4_K_M— context 1024k tokens - qwen3.8-flash-next
IQ4_XS— context 256k tokens - deepseek-v4-flash
IQ4_XS— context 1024k tokens - glm-5.2
Q4_K_M— context 16k tokens
Models that do not fit — 0 of the 20 models measured
All 20 measured models fit on this machine.
The machines on either side
- Xiaomi AI Cube, 80 GB — engineering prototype — 77.00 GiB left for the model, 17 / 20 models
- PC with Ryzen AI Max+ 395 and 128 GB unified — 96.00 GiB left for the model, 18 / 20 models
- Mac Studio with M5 Max and 128 GB — 125.00 GiB left for the model, 18 / 20 models
- PC with Ryzen AI Max+ PRO 495 and 192 GB unified — 160.00 GiB left for the model, 19 / 20 models
Price and availability
What a machine costs decides as much as what fits inside it, so the price belongs on this page. What follows is only what the manufacturer publishes, with the market it applies to and the day it was read.
From $5,499 (United States), as published on 2026-09-23 (Source). Vendor declared
That is the starting price of the Mac Studio with M5 Ultra, which is the 96 GB base — not the 512 GB this sheet is computed from. Apple’s store does not price that configuration, and we can say where we looked: read on 23 September 2026, its memory selector gives late October for the 512 GB option, and the URL that would hold its price does not exist. No price in euros either: this figure is United States dollars.
What we earn on this page
Nothing, today. There is no affiliate link on this page and no commercial agreement behind any figure on it. If a buying link ever appears here, it will be marked as such and the commission declared in this same block. Two rules do not change on that day: the ranking on this site is computed before commercial availability is applied, and «best» never means «pays most».



