Hot on the heels of the new Mac mini announcement, Apple today (August 25) unveiled a new generation of Mac Studio, built for the demands of serious professionals and AI development workflows. Retaining the model’s signature compact desktop design, this Mac Studio skips straight to the M5 Max and debuts the most powerful quad-die architecture Apple Silicon has ever produced: the M5 Ultra. The lineup once again supports linking multiple units together via Thunderbolt 5, fundamentally redefining the ceiling of local edge AI computing.
Breaking the Compute Ceiling: M5 Max and M5 Ultra Arrive Together
The core highlight of the new Mac Studio lies in its comprehensive adoption of Neural Accelerators within the GPU architecture, paired with an astonishingly large unified memory capacity, positioning this desktop as the most powerful locally hosted large language model inference engine currently on the market.
M5 Max: The Performance Foundation for High-End Professionals
- Core Configuration: An 18-core CPU (6 performance cores plus 12 efficiency cores) paired with up to a 40-core GPU.
- AI Performance Leap: Built-in Neural Accelerators within the GPU deliver up to a 3.9x improvement in AI compute performance over the previous generation. Supports up to 128GB of unified memory with 614GB/s of memory bandwidth.
M5 Ultra: The Ultimate Quad-Die Monster
- Core Configuration: Up to a 36-core CPU paired with an 80-core GPU.
- AI and Video Editing at the Extreme: Peak AI compute performance reaches 4.3 times that of the M3 Ultra and 9.8 times that of the original M1 Ultra. The formidable media engine can smoothly play back as many as 33 simultaneous streams of 8K ProRes 422 video.
- 512GB of Massive Memory: Supports an industry-rare configuration of up to 512GB of unified memory, with bandwidth soaring to 1.2TB/s – meaning data scientists can load enormous open-weight models entirely into local memory to run them, completely eliminating dependence on cloud APIs and the privacy concerns that come with them.
Key insight: through its UltraFusion quad-die architecture, the M5 Ultra roughly doubles the M5 Max’s raw physical hardware specifications, and its 512GB unified memory configuration in particular delivers an absolute advantage when running exceptionally high-parameter large language models.
Thunderbolt 5 Debuts: RDMA Support Enables a Desktop AI Cluster
On the I/O front, the new Mac Studio introduces Thunderbolt 5 for the first time, delivering an astonishing 120Gb/s of bandwidth and supporting hardware clustering over the connection.
Through Thunderbolt 5’s built-in RDMA (Remote Direct Memory Access) capability, users can link multiple Mac Studio units together to create an enormously large shared memory pool. According to Apple’s official product page, a compute cluster built from four Mac Studio units can deliver up to 3 times the distributed AI inference performance of a single system – offering enterprises that need to deploy massive language models on-premises a highly flexible and cost-effective alternative to a traditional data center buildout.
The new machine also introduces support for Wi-Fi 7 and Bluetooth 6.0 for the first time, and can drive up to 8 external displays (or 4 Studio Display XDR panels running at 5K and 120Hz).
Availability and Pricing
The new Mac Studio is available for preorder starting today across 30 countries, including the United States, with shipping expected to begin September 22 (the top-tier configuration with 512GB of memory is expected to ship in late October).
- M5 Max configuration: Starting at $2,499
- M5 Ultra configuration: Starting at $5,499
No Servers Needed: Apple Reimagines the Edge AI Data Center With Mac Studio
The biggest surprise Apple brought to this Mac Studio refresh isn’t simply stacking more silicon – it’s opening up “multi-unit clustering” via Thunderbolt 5 and RDMA technology. A capability once reserved for data-center-grade high-end networking architectures, such as NVIDIA’s NVLink or InfiniBand, has now trickled down to desktop-class hardware.
At a moment when NVIDIA dominates cloud compute and the high-performance server market almost entirely, Apple has chosen a different path: leveraging the M5 Ultra’s 512GB of unified memory and multi-unit clustering capability to build enterprises and research institutions a desktop-class AI data center that requires no elaborate cooling infrastructure and runs remarkably quiet and power-efficient. When developers can load an entire open-source model – such as Llama – directly into local memory to train and run inference, it clearly opens up an application model fundamentally different from building AI systems atop public cloud platforms like Microsoft’s.
Support Our Threat Intelligence
If you find our technology report and cybersecurity news helpful, consider supporting our work.