Mac Studio Ultra
Apple Silicon, unified memory

| Indicative price | €6,599 to €20,679 incl. VAT on the Apple Store France as of 25 August 2026 (as reported by MacGeneration): €6,599 for the M5 Ultra with 96 GB and 1 TB, €20,679 with 36 cores, 256 GB and 16 TB; the 512 GB variant, announced for late October 2026, is not priced yet (price guide) |
|---|---|
| Memory | M5 Ultra up to 512 GB: 96 GB base, 256 or 512 GB as options (36-core CPU variant only, 512 GB announced for late October 2026) |
| Chip | M5 Ultra, 30 CPU cores and 64 GPU cores, or 36 and 80 as an option (detailed analysis) |
| Bandwidth | 1.2 TB/s unified, 50% more than the M3 Ultra according to Apple (819 GB/s on the 2025 datasheet, i.e. +47% by calculation) |
| Runtime | MLX or llama.cpp |
| Power draw | 480 W maximum continuous (Apple datasheet), quiet |
| Form factor | Compact desktop |
What the machine runs
Large quantised mixture-of-experts models thanks to unified memory, with moderate throughput. Recommended runtime: MLX or llama.cpp.
Positioning
Apple's option remains quiet and offers large memory per euro and per watt. The 2026 RAM shortage means availability and pricing must be checked.
Manufacturers
The Mac Studio is designed and assembled by Apple.
Support
The Mac Studio runs on Apple Silicon. NVIDIA AI Enterprise does not apply to this hardware; support comes from QDNA and the open-source ecosystem, MLX and llama.cpp.
Which models fit on Mac Studio Ultra?
The machine offers 512 GB of memory. 4 of the 7 open models in the catalogue load on it, in the format shown.
| Model | Billion parameters | Most precise format that fits | Sizing page |
|---|---|---|---|
| GLM 5.2 | 744 | NVFP4 | GLM 5.2 on Mac Studio Ultra |
| Nemotron 3 Ultra | 550 | NVFP4 | Nemotron 3 Ultra on Mac Studio Ultra |
| MiniMax M3 | 428 | FP8 | MiniMax M3 on Mac Studio Ultra |
| Qwen 3.8 27B | 27 | FP16 | Qwen 3.8 27B on Mac Studio Ultra |
See the full sizing matrix.
Frequently asked questions
How much does a configured mac studio cost for an LLM?
Pricing depends on configuration and supplier. The page shows a dated public price; we provide a detailed quote after a scoping call.
Which LLM fits in a mac studio?
Depends on the quantization format and the input context. The table on the page indicates the recommended maximum model size.
Which runtimes support the mac studio?
MLX or llama.cpp; vLLM and Triton Inference Server target NVIDIA GPUs. The model, format and GPU combinations measured by QDNA are published in the measurements section.