> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Apple M5 Ultra Nearly Quadruples DGX Spark Token Speed in Qwen Test
- URL: https://www.implicator.ai/apple-m5-ultra-dgx-spark-qwen-test/
- Published: 2026-09-21T15:53:45.000Z
- Updated: 2026-09-21T15:53:45.000Z
- Description: Apple’s M5 Ultra Mac Studio generated tokens nearly four times as fast as Nvidia’s DGX Spark in a test of one Qwen model. The tested configuration cost $12,299, with a $4,000 charge to move from 96GB to 256GB of memory. Results varied across other workloads, and the benchmark does not establish end-
- Author: Marcus Schuler
- Tags: Tools & Workflows

Apple’s [M5 Ultra Mac Studio](https://www.apple.com/mac-studio/?ref=implicator.ai) generated tokens nearly four times as fast as Nvidia’s DGX Spark in a review benchmark of a single Qwen model. The tested configuration cost $12,299, so the benchmark’s performance gain comes with a five-figure purchase price.

In a [review published September 21](https://www.tomshardware.com/desktops/mini-pcs/apple-mac-studio-m5-ultra-review?ref=implicator.ai), Andrew E. Freedman tested Qwen 3.8-27B-Q4\_K\_M, a dense, quantized model whose performance depends heavily on memory bandwidth. Prompt processing measures how quickly a system processes the input context before generation. Token generation measures how quickly it produces output tokens.

Because this was one dense-model test aimed largely at GPU performance in isolation, the result does not establish end-to-end agent performance across other models, context lengths or software stacks.

What Changed

- M5 Ultra generated tokens nearly four times as fast as DGX Spark in a review benchmark of Qwen 3.8-27B-Q4\_K\_M.
- The tested 256GB Mac Studio cost $12,299; the $5,499 base uses a lower-core chip.
- Upgrading Ultra memory from 96GB to 256GB costs $4,000, with no later RAM upgrade.
- Other workloads showed smaller gains, and one HandBrake test finished slower than the M3 Ultra.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## A bandwidth-heavy test

The comparison also showed twice the generation throughput of an M4 Max Mac and faster prompt processing than DGX Spark. The tested Mac Studio paired a 36-core CPU and 80-core GPU with 256GB of unified memory and a 4TB SSD. The top chip adds $1,300 to the M5 Ultra base version, which has a 30-core CPU and 64-core GPU. That chip charge is separate from memory and storage upgrades.

FREE WEEKDAY MORNING BRIEFING

Follow the tests that matter for local AI.

The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

Email address 

Send me tomorrow’s briefing 

Check your inbox for the confirmation link.

About five minutes. No hype. No spam.

The unified-memory design gives the CPU and GPU access to a shared memory pool. The 256GB capacity lets the machine hold very large individual models or several models on one node. Apple lists memory bandwidth of 1.2 terabytes per second for the M5 Ultra, compared with up to 614 gigabytes per second for the M5 Max. The M5 Max supports up to 128GB of unified memory.

## Memory carries a steep premium

The M5 Ultra Mac Studio starts at $5,499, compared with $2,499 for the M5 Max version. Buyers who need 256GB must move to an M5 Ultra configuration. Moving from the Ultra’s 96GB base to 256GB adds $4,000, and buyers cannot install more memory later.

The M3 Ultra Mac Studio launched at $3,999 in 2025\. Apple raised its starting price to [$5,299 in June 2026](https://www.macworld.com/article/3238319/mac-studio-m5-max-review.html?ref=implicator.ai), before setting the new M5 Ultra base at $5,499\. The latest increment is $200\. The $1,500 launch-to-launch difference uses the M3 Ultra’s original price.

Steve Dent concluded that the Ultra’s premium over the M5 Max was difficult to justify outside demanding professional work. In his [synthetic Geekbench AI tests](https://www.engadget.com/2263184/apple-mac-studio-m5-ultra-review/?ref=implicator.ai), the M5 Ultra was about 10% faster than the M5 Max. That measurement covers a different workload from Freedman’s Qwen inference test and is not a replication of it.

Know someone who'd find this useful? [✉️ Email it to a friend in one click](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20maybe%20three%20newsletters%20I%20actually%20read.%20The%20rest%20just%20pile%20up%2C%20unread%2C%20judging%20me.%0A%0AAnd%20yes%2C%20this%20email%20mostly%20wrote%20itself%2C%20which%20is%20a%20little%20on%20the%20nose%20for%20an%20AI%20newsletter.%20Doesn%27t%20make%20it%20wrong.%20implicator.ai%20is%20good.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dforward%26utm%5Fcampaign%3Demail%5Fforward), or they can [subscribe free here](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=forward&utm%5Fcampaign=forward%5Fto%5Fcolleague).

## Results by workload

Lori Grunin measured a [72% higher score](https://www.cnet.com/tech/computing/apple-mac-studio-m5-ultra-review/?ref=implicator.ai) for the 80-core M5 Ultra than the 40-core M5 Max in Procyon’s Stable Diffusion XL test, while the Ultra’s GPU scored 46% higher than the M5 Max in Cinebench 2026 rendering. Brian Westover’s [HandBrake conversion](https://www.pcmag.com/reviews/apple-mac-studio-2026-m5-ultra?ref=implicator.ai) took 87 seconds, three seconds longer than the M3 Ultra system in his comparison.

Unlike Nvidia and AMD, which also offer desktops capable of local AI work, Apple offered no guides or playbooks for these uses, including for the developer community, Westover observed. Buyers would need to work out those uses themselves. Software updates and compatibility conflicts also prevented Westover from gathering reliable Blender and Premiere comparison data.

## What ships next

Apple lists the new Mac Studio for general availability on September 22\. As of his review, Freedman observed that the tested $12,299 configuration was backordered between 16 and 18 weeks. That lead time applied to the tested configuration, not every Mac Studio model. Configurations with 512GB of unified memory are scheduled to follow in late October.

Frequently Asked Questions

Does the M5 Ultra run all AI tasks four times faster?

No. The comparison used one dense Qwen model and largely isolated GPU performance. It does not establish end-to-end agent performance across models, contexts or software stacks.

Which Mac Studio was tested?

The review unit combined a 36-core CPU, 80-core GPU, 256GB of unified memory and a 4TB SSD for $12,299.

Does the $5,499 base model have the same chip?

No. The base M5 Ultra has a 30-core CPU and 64-core GPU. The tested 36-core/80-core chip adds $1,300 before memory and storage upgrades.

How much does the extra memory cost?

Moving from 96GB to 256GB costs $4,000\. Memory cannot be upgraded after purchase. The M5 Max supports up to 128GB.

When are the new configurations available?

Apple lists September 22 availability. The tested configuration was backordered 16-18 weeks as of the review. The 512GB option is scheduled for late October.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## Related stories

[Apple Mac Studio M5 Ultra Targets Local AI Prompt SpeedApple's new Mac Studio adds matrix compute to every GPU core and scales to 512GB of unified memory. The gain lands on prompt processing, not on how fast tokens stream back. At $2,499 and $5,499 these are the line's highest starting prices.Implicator.ai![](https://www.implicator.ai/content/images/2026/08/20260825-224728-waiting_first_token_v2.webp)](https://www.implicator.ai/apple-mac-studio-m5-ultra-prompt-processing/)

[Perplexity Ships Local AI Agent on Nvidia DGX SparkPerplexity moved its entire agent stack onto hardware users own, starting with Nvidia's DGX Spark. Work done on the device consumes no billing credits, and every cloud escalation needs separate approval. Security practitioners say a permission prompt is consent, not control.Implicator.ai![](https://www.implicator.ai/content/images/2026/08/20260826-034018-analyst_local_box.webp)](https://www.implicator.ai/perplexity-ships-local-ai-agent-on-nvidia-dgx-spark-with-no-on-device-token-cost/)

[Apple Weighs M8 Ultra Servers With Nvidia NVLink FusionApple is weighing a return to selling servers, a rack-mounted line built on two or four future M8 Ultra chips, and has discussed using Nvidia's NVLink Fusion. Nothing is planned before 2029, and buyers would include AI developers, businesses and governments.Implicator.ai![](https://www.implicator.ai/content/images/2026/09/2026-09-16-apple-m8-ultra-servers-nvlink-fusion.webp)](https://www.implicator.ai/apple-m8-ultra-servers-nvlink-fusion/)