Apple’s M5 Ultra Mac Studio generated tokens nearly four times as fast as Nvidia’s DGX Spark in a review benchmark of a single Qwen model. The tested configuration cost $12,299, so the benchmark’s performance gain comes with a five-figure purchase price.

In a review published September 21, Andrew E. Freedman tested Qwen 3.8-27B-Q4_K_M, a dense, quantized model whose performance depends heavily on memory bandwidth. Prompt processing measures how quickly a system processes the input context before generation. Token generation measures how quickly it produces output tokens.

Because this was one dense-model test aimed largely at GPU performance in isolation, the result does not establish end-to-end agent performance across other models, context lengths or software stacks.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

A bandwidth-heavy test

The comparison also showed twice the generation throughput of an M4 Max Mac and faster prompt processing than DGX Spark. The tested Mac Studio paired a 36-core CPU and 80-core GPU with 256GB of unified memory and a 4TB SSD. The top chip adds $1,300 to the M5 Ultra base version, which has a 30-core CPU and 64-core GPU. That chip charge is separate from memory and storage upgrades.

The unified-memory design gives the CPU and GPU access to a shared memory pool. The 256GB capacity lets the machine hold very large individual models or several models on one node. Apple lists memory bandwidth of 1.2 terabytes per second for the M5 Ultra, compared with up to 614 gigabytes per second for the M5 Max. The M5 Max supports up to 128GB of unified memory.

Memory carries a steep premium

The M5 Ultra Mac Studio starts at $5,499, compared with $2,499 for the M5 Max version. Buyers who need 256GB must move to an M5 Ultra configuration. Moving from the Ultra’s 96GB base to 256GB adds $4,000, and buyers cannot install more memory later.

The M3 Ultra Mac Studio launched at $3,999 in 2025. Apple raised its starting price to $5,299 in June 2026, before setting the new M5 Ultra base at $5,499. The latest increment is $200. The $1,500 launch-to-launch difference uses the M3 Ultra’s original price.

Steve Dent concluded that the Ultra’s premium over the M5 Max was difficult to justify outside demanding professional work. In his synthetic Geekbench AI tests, the M5 Ultra was about 10% faster than the M5 Max. That measurement covers a different workload from Freedman’s Qwen inference test and is not a replication of it.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Results by workload

Lori Grunin measured a 72% higher score for the 80-core M5 Ultra than the 40-core M5 Max in Procyon’s Stable Diffusion XL test, while the Ultra’s GPU scored 46% higher than the M5 Max in Cinebench 2026 rendering. Brian Westover’s HandBrake conversion took 87 seconds, three seconds longer than the M3 Ultra system in his comparison.

Unlike Nvidia and AMD, which also offer desktops capable of local AI work, Apple offered no guides or playbooks for these uses, including for the developer community, Westover observed. Buyers would need to work out those uses themselves. Software updates and compatibility conflicts also prevented Westover from gathering reliable Blender and Premiere comparison data.

What ships next

Apple lists the new Mac Studio for general availability on September 22. As of his review, Freedman observed that the tested $12,299 configuration was backordered between 16 and 18 weeks. That lead time applied to the tested configuration, not every Mac Studio model. Configurations with 512GB of unified memory are scheduled to follow in late October.

Frequently Asked Questions

Does the M5 Ultra run all AI tasks four times faster?

No. The comparison used one dense Qwen model and largely isolated GPU performance. It does not establish end-to-end agent performance across models, contexts or software stacks.

Which Mac Studio was tested?

The review unit combined a 36-core CPU, 80-core GPU, 256GB of unified memory and a 4TB SSD for $12,299.

Does the $5,499 base model have the same chip?

No. The base M5 Ultra has a 30-core CPU and 64-core GPU. The tested 36-core/80-core chip adds $1,300 before memory and storage upgrades.

How much does the extra memory cost?

Moving from 96GB to 256GB costs $4,000. Memory cannot be upgraded after purchase. The M5 Max supports up to 128GB.

When are the new configurations available?

Apple lists September 22 availability. The tested configuration was backordered 16-18 weeks as of the review. The 512GB option is scheduled for late October.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Apple Mac Studio M5 Ultra Targets Local AI Prompt Speed
Apple's new Mac Studio adds matrix compute to every GPU core and scales to 512GB of unified memory. The gain lands on prompt processing, not on how fast tokens stream back. At $2,499 and $5,499 these are the line's highest starting prices.
Perplexity Ships Local AI Agent on Nvidia DGX Spark
Perplexity moved its entire agent stack onto hardware users own, starting with Nvidia's DGX Spark. Work done on the device consumes no billing credits, and every cloud escalation needs separate approval. Security practitioners say a permission prompt is consent, not control.
Apple Weighs M8 Ultra Servers With Nvidia NVLink Fusion
Apple is weighing a return to selling servers, a rack-mounted line built on two or four future M8 Ultra chips, and has discussed using Nvidia's NVLink Fusion. Nothing is planned before 2029, and buyers would include AI developers, businesses and governments.
Tools & Workflows

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai