DGX Spark vs. Cloud GPUs: Buy for Memory, Rent First

NVIDIA's DGX Spark priced and dated against live cloud GPU rates: what its 128GB memory can run, its bandwidth limit, and when to rent instead.

By Faiz Ahmed•
•12 min read

The NVIDIA DGX Spark is a desktop AI workstation built around the GB10 Grace Blackwell superchip, with 128 GB of unified memory shared between its Arm CPU and GPU, per NVIDIA's technical documentation. StorageReview reported the Founders Edition (4TB) system's $3,999 launch price on 14 October 2025, and Tom's Hardware reported on 27 February 2026 that memory shortages pushed the Founders Edition system price up by $700 to $4,699, a derived 17.5 percent increase. The rule for choosing between a Spark and a rented GPU is simple: buy the Spark for local development when its shared memory suits your model, and rent first for training runs or short projects before committing to hardware.

GB10 pairs a 20-core Arm CPU, ten Cortex-X925 cores and ten Cortex-A725 cores, with a Blackwell GPU, per NVIDIA's technical documentation. That documentation, checked 28 September 2026, describes the 128 GB of LPDDR5x as "coherent unified system memory" shared between the two, rather than separate CPU and GPU pools, with 1 TB or 4 TB self-encrypting NVMe storage options for DGX Spark. NVIDIA's product page, also checked 28 September 2026, gives a 5.9 by 5.9 by 2 inch case, about 1.1 litres, a GB10 chip thermal design power of 140 W, and a 240 W power supply. It also lists an onboard ConnectX-7 SmartNIC and 10 Gigabit Ethernet, and says NVIDIA's AI software stack ships preinstalled.

What a DGX Spark has cost

The price has changed twice since NVIDIA first named the product.

MilestonePriceDateSource
Announced as Project DIGITS$3,000 (system starting price)6 January 2025NVIDIA
Founders Edition (4TB), reported system launch price$3,99914 October 2025StorageReview
Founders Edition system, reported price after the memory-shortage increase$4,69927 February 2026Tom's Hardware

StorageReview's independent review, published on 14 October 2025, confirmed the Founders Edition (4TB) system's launch price at $3,999. The later increase Tom's Hardware reported lines up with the wider DRAM market: the same outlet reported on 23 August 2026, citing people familiar with the matter rather than NVIDIA confirmation, that NVIDIA had warned some of its largest customers of server price rises exceeding 15 percent in many cases from early 2027. It cited analyst estimates projecting conventional DRAM contract prices rising 58 to 63 percent quarter over quarter in the second quarter of 2026. Samsung and SK hynix separately raised their 2026 HBM3E supply prices by close to 20 percent, per a TrendForce report dated 24 December 2025. This is reported market context, not a breakdown of the Spark's costs or NVIDIA's margin.

NVIDIA's launch release, dated 13 October 2025, names Acer, ASUS, Dell Technologies, GIGABYTE, HP, Lenovo, and MSI as partners building GB10 systems. Arm's newsroom, dated 23 March 2026, names the specific products: the Acer Veriton GN100, ASUS Ascent GX10, Dell Pro Max with GB10, Gigabyte AI TOP ATOM, HP ZGX Nano AI Station, Lenovo ThinkStation PGX, and MSI EdgeXpert, each built around "unified memory architectures of up to 128GB." Get a current quote from the seller before comparing one with the Founders Edition.

What NVIDIA claims, and what reviewers measured

NVIDIA's claims, by date

NVIDIA's own claims about DGX Spark's model-size ceiling describe different configurations and tasks. On 6 January 2025, unveiling the device as "Project DIGITS," NVIDIA said developers could run "up to 200-billion-parameter large language models" on one unit, and that "two Project DIGITS AI supercomputers can be linked to run up to 405-billion-parameter models," both vendor claims. In its launch release on 13 October 2025, NVIDIA described a single unit running "inference on AI models with up to 200 billion parameters" and fine-tuning "models of up to 70 billion parameters locally," again vendor claims, and it did not repeat the two-unit 405B figure. NVIDIA's product page, checked 28 September 2026, describes a larger configuration: it says ConnectX networking "enables the connection of up to four NVIDIA DGX Spark systems to work with AI models of up to 700 billion parameters," another vendor claim. None of these model-size limits is an independent measurement. Treat each as what NVIDIA said its hardware could do on the date it said it.

NVIDIA's technical documentation lists FP4 compute of "up to 1,000 TOPS" for inference and "up to 1 PFLOP" at FP4 precision with sparsity. Its product page, checked 28 September 2026, footnotes that PFLOP figure as "Theoretical FP4 TOPS using the sparsity feature," confirming it is the sparse number, not the dense one. Treat this vendor claim as a theoretical sparse peak, not measured application throughput.

What it runs, measured

The measurements below were published on 13 and 14 October 2025. Treat them as dated results whose software and setup may differ from yours, not as measurements of today's performance.

LMSYS, publishing through the SGLang team on 13 October 2025, reported gpt-oss-20b results in MXFP4 precision through Ollama on three systems.

SystemPrefill (tokens/sec)Decode (tokens/sec)
DGX Spark2,05349.7
RTX PRO 6000 Blackwell Workstation Edition10,108215
RTX 50908,519205

On this one task, the RTX PRO 6000 decoded about 4.3 times faster than the Spark and the RTX 5090 about 4.1 times faster, both derived from the figures above.

The same LMSYS post shows why batch size matters. Running the larger Llama 3.1 70B in FP8 through SGLang, DGX Spark measured "803 tps prefill / 2.7 tps decode." Running the smaller Llama 3.1 8B in FP8, also through SGLang, tells the other half of the story.

Batch sizePrefill (tokens/sec)Decode (tokens/sec)
17,99120.5
327,949368

Prefill barely moved, but decode rose about 18.0 times from batch 1 to batch 32, derived. LMSYS's own conclusion matches that result: the unified LPDDR5x memory, "offering up to 273 GB/s, shared across both CPU and GPU," is "the key bottleneck in AI inference performance." Across LMSYS, ServeTheHome and StorageReview, the reported bandwidth limitation affects single-stream generation on large dense models more than prefill, batched or mixture-of-experts workloads.

ServeTheHome's 14 October 2025 review, running Ollama out of the box, measured gpt-oss-20b "often over 49 tokens/second," gpt-oss-120b at "14.48 tokens/second," and Qwen3 32B at "9-10 tokens/second." StorageReview, reviewing the device the same day, measured Llama 3.1 8B decode rising from "13.6 tok/s" at BF16 to "23.2" at FP8 to "34.1" at FP4 at concurrency 1, with reported higher-concurrency results of "408.6," "752.8" and "924.1" tok/s respectively. It explicitly gives concurrency 128 for the BF16 result. It put gpt-oss-120b in NVFP4 at "31.4 tok/s at concurrency 1 and 162.7 tok/s at 64 concurrent requests," and measured achieved matrix-multiply throughput of "99.8 TFLOPs" at BF16 and "207.7 TFLOPs" at FP8, both measured rather than theoretical peaks. These results use different setups; do not treat them as a controlled comparison between reviewers.

Reported limitations

Independent reviewers flagged three limits worth knowing before you buy. LMSYS's own read, quoted above, is that 273 GB/s of shared memory bandwidth is a bottleneck for single-stream generation on large dense models. Simon Willison, writing on 15 October 2025, ran into early ARM64 software friction: obtaining PyTorch wheels built for Arm CUDA was difficult, NVIDIA's documentation "changed substantially" within a week of launch, and he judged "It's a bit too early for me to provide a confident recommendation." ServeTheHome, on 14 October 2025, found the dual-port networking undershooting its own spec: combined throughput across both QSFP ports "maxed out at approximately 100Gbps," about half the advertised figure.

Memory bandwidth: the real speed limit

NVIDIA's developer blog, dated 17 November 2023, explains why memory bandwidth can set the pace of text generation. During decoding, it states, "The speed at which the data (weights, keys, values, activations) is transferred to the GPU from memory dominates the latency, not how fast the computation actually happens. In other words, this is a memory-bound operation." The same post describes reading your prompt, the step called prefill, as a highly parallel matrix operation that saturates GPU utilization. This explains the bandwidth comparison below; it is not a runtime prediction for every model or batch size.

DGX Spark's own memory bandwidth is 273 GB/s, per NVIDIA's technical documentation. The site's spec table below shows the corresponding specifications for three GPU families to compare with it.

SpecRTX 5090RTX PRO 6000H100
VRAM32 GB96 GB80 to 94 GB
Memory typeGDDR7GDDR7HBM3
Memory bandwidth1,792 GB/s1,792 GB/s3,350 GB/s
FP4 (dense)1,676 TFLOPS2,015.2 TFLOPSNot published
FP4 (with sparsity)3,352 TFLOPS4,030.4 TFLOPSNot published
TDP575 W600 W700 W
Figures from the vendor datasheets: RTX 5090, RTX PRO 6000, H100, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

By that table, the RTX 5090 and the RTX PRO 6000 Blackwell each carry 1,792 GB/s of memory bandwidth, about 6.6 times the Spark's 273 GB/s (derived). The H100 family entry gives 3,350 GB/s, about 12.3 times the Spark's figure (derived). Apple's Mac Studio provides another memory-bandwidth comparison. Apple's specification page, checked 28 September 2026, lists the base M5 Max configuration at 460 GB/s, about 1.7 times the Spark's bandwidth (derived), and the same chip configured with a 40-core GPU at 614 GB/s, about 2.2 times (derived). It lists the M5 Ultra at 1.2 TB/s, or 1,200 GB/s, about 4.4 times the Spark's bandwidth (derived). These are specification ratios, not measured application speedups.

The trade-off is now clear: NVIDIA's documented 128 GB of shared system memory exceeds the dedicated GPU memory in each family in the table above, and the base Mac Studio memory capacities Apple lists. Shared system memory is not the same as dedicated GPU memory, so check your model's requirements rather than treating the whole pool as model storage. The Spark's listed bandwidth is the lowest of this comparison group, and LMSYS's gpt-oss-20b results above show slower decode than the two RTX cards on that task. As local LLM hardware, the Spark offers a memory-capacity trade-off, not a general speed advantage. For a longer walkthrough of reading a spec sheet this way, see how to read GPU specs for AI, and check exact model memory needs with the LLM VRAM calculator before you decide.

What that budget buys in rented GPU-hours

At the $4,699 Founders Edition price Tom's Hardware reported on 27 February 2026, this is what the same money buys in rented GPU time. Check NVIDIA's store for today's price before you decide.

GPU$/GPU-hrProviderGPU-hours for $4,699Days at 24 h a day
H100$2.50Hyperstack1,87978.3
RTX PRO 6000 Blackwell$0.59RunPod7,964331.8
RTX 5090$0.53Vast.ai8,866369.4
RTX 4090$0.43Vast.ai10,927455.3
A100$0.68LeaderGPU6,910287.9
What $4,699 buys at the cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. GPU-hours = $4,699 ÷ hourly price, rounded down; days = GPU-hours ÷ 24. Storage, data transfer and tax are not included. Latest stock observation: .

Read the table above as an opportunity cost, not a verdict. It converts the purchase budget into GPU-hours; it does not establish how much work your application completes in that time. Before buying, compare your measured workload, schedule and full costs in the rent versus buy GPU calculator. Check the rental's billing and uptime terms rather than assuming every offer works the same way.

Buy the Spark, or rent: five situations

The right answer changes with the job, not just the budget.

SituationBuy Spark or rentWhy
Prototype agents locallyBuy SparkChoose local development if NVIDIA's documented shared memory suits the model and context you need. Configure the workflow to keep your data on site.
Fine-tune a 70B modelRent firstNVIDIA's 13 October 2025 vendor claim includes 70B fine-tuning, but no published review times that run against an H100 or RTX PRO 6000 Blackwell. Compare your actual run before buying.
Serve a model to usersRent firstEvaluate throughput at your target concurrency and the provider's uptime terms before committing to production.
Run overnight batch jobsBuy Spark if speed can waitSchedule the job and check the results in the morning, but include electricity in your cost comparison.
Data that cannot leave the buildingBuy SparkKeep the workload on your own network and check that its software and integrations meet that requirement.

The decision rule

Buy a DGX Spark when the bottleneck is local memory capacity and you can live with its bandwidth: local prototyping, models you only need to hold and query occasionally, and workloads you need to keep in the building. Rent first for training, fine-tuning, and serving real traffic, then measure your own workload before buying. Use the published RTX results above as evidence for those specific tests, not as an H100 or RTX 5090 guarantee for every job. The DGX systems explainer is the place to compare other DGX systems, and LLM GPU requirements is worth checking for the model you actually plan to run before you commit to either path.

If you want to try an LLM before buying hardware, rent first to learn what you need. If you already know you need a local AI workstation and the Spark's shared memory and measured performance suit your workload, buy the Spark.

Sources

Frequently asked questions

What is the NVIDIA DGX Spark?▾

NVIDIA describes DGX Spark as a desktop AI system built around its GB10 Grace Blackwell superchip, with 128 GB of unified memory shared between its Arm CPU and GPU. NVIDIA sells a Founders Edition, and seven partners, including ASUS, Dell, HP and Lenovo, sell their own GB10 systems.

How much does a DGX Spark cost?▾

StorageReview reported a $3,999 launch price for the Founders Edition (4TB) system on 14 October 2025; NVIDIA had announced Project DIGITS at a $3,000 system starting price on 6 January 2025. Tom's Hardware reported on 27 February 2026 that memory shortages pushed the Founders Edition system price up by $700 to $4,699, a derived 17.5 percent increase.

Can a DGX Spark fine-tune a 70 billion parameter model?▾

NVIDIA's 13 October 2025 release makes that claim for a single unit, but it is a vendor claim, not an independent measurement. The cited sources do not establish how long that fine-tuning run takes on a Spark versus an H100 or RTX PRO 6000 Blackwell.

Is a DGX Spark faster than an RTX 5090?▾

In LMSYS's 13 October 2025 gpt-oss-20b MXFP4 test through Ollama, the RTX 5090 decoded about 4.1 times faster than the Spark, derived from the reported results. NVIDIA's Spark documentation and the site's GPU spec table show more shared memory capacity on the Spark but lower memory bandwidth; that does not establish a speed ranking for every workload.

Should I buy a DGX Spark or rent a cloud GPU?▾

Buy a Spark for local development when its shared memory suits your model and keeping the workload on site matters. Rent first for training, fine-tuning, or serving traffic, and compare the workload's measured speed before buying.

How does a DGX Spark compare to a Mac Studio?▾

Apple's Mac Studio specification page, checked 28 September 2026, lists 1.2 TB/s of memory bandwidth for the M5 Ultra, above the Spark's 273 GB/s in NVIDIA's documentation. NVIDIA says the Spark ships with its AI stack preinstalled; the bandwidth comparison alone does not establish application performance.

Related Posts