The NVIDIA DGX Spark is a desktop AI workstation built around the GB10 Grace Blackwell superchip, with 128 GB of unified memory shared between its Arm CPU and GPU, per NVIDIA's technical documentation. StorageReview reported the Founders Edition (4TB) system's $3,999 launch price on 14 October 2025, and Tom's Hardware reported on 27 February 2026 that memory shortages pushed the Founders Edition system price up by $700 to $4,699, a derived 17.5 percent increase. The rule for choosing between a Spark and a rented GPU is simple: buy the Spark for local development when its shared memory suits your model, and rent first for training runs or short projects before committing to hardware.
GB10 pairs a 20-core Arm CPU, ten Cortex-X925 cores and ten Cortex-A725 cores, with a Blackwell GPU, per NVIDIA's technical documentation. That documentation, checked 28 September 2026, describes the 128 GB of LPDDR5x as "coherent unified system memory" shared between the two, rather than separate CPU and GPU pools, with 1 TB or 4 TB self-encrypting NVMe storage options for DGX Spark. NVIDIA's product page, also checked 28 September 2026, gives a 5.9 by 5.9 by 2 inch case, about 1.1 litres, a GB10 chip thermal design power of 140 W, and a 240 W power supply. It also lists an onboard ConnectX-7 SmartNIC and 10 Gigabit Ethernet, and says NVIDIA's AI software stack ships preinstalled.
What a DGX Spark has cost
The price has changed twice since NVIDIA first named the product.
| Milestone | Price | Date | Source |
|---|---|---|---|
| Announced as Project DIGITS | $3,000 (system starting price) | 6 January 2025 | NVIDIA |
| Founders Edition (4TB), reported system launch price | $3,999 | 14 October 2025 | StorageReview |
| Founders Edition system, reported price after the memory-shortage increase | $4,699 | 27 February 2026 | Tom's Hardware |
StorageReview's independent review, published on 14 October 2025, confirmed the Founders Edition (4TB) system's launch price at $3,999. The later increase Tom's Hardware reported lines up with the wider DRAM market: the same outlet reported on 23 August 2026, citing people familiar with the matter rather than NVIDIA confirmation, that NVIDIA had warned some of its largest customers of server price rises exceeding 15 percent in many cases from early 2027. It cited analyst estimates projecting conventional DRAM contract prices rising 58 to 63 percent quarter over quarter in the second quarter of 2026. Samsung and SK hynix separately raised their 2026 HBM3E supply prices by close to 20 percent, per a TrendForce report dated 24 December 2025. This is reported market context, not a breakdown of the Spark's costs or NVIDIA's margin.
NVIDIA's launch release, dated 13 October 2025, names Acer, ASUS, Dell Technologies, GIGABYTE, HP, Lenovo, and MSI as partners building GB10 systems. Arm's newsroom, dated 23 March 2026, names the specific products: the Acer Veriton GN100, ASUS Ascent GX10, Dell Pro Max with GB10, Gigabyte AI TOP ATOM, HP ZGX Nano AI Station, Lenovo ThinkStation PGX, and MSI EdgeXpert, each built around "unified memory architectures of up to 128GB." Get a current quote from the seller before comparing one with the Founders Edition.
What NVIDIA claims, and what reviewers measured
NVIDIA's claims, by date
NVIDIA's own claims about DGX Spark's model-size ceiling describe different configurations and tasks. On 6 January 2025, unveiling the device as "Project DIGITS," NVIDIA said developers could run "up to 200-billion-parameter large language models" on one unit, and that "two Project DIGITS AI supercomputers can be linked to run up to 405-billion-parameter models," both vendor claims. In its launch release on 13 October 2025, NVIDIA described a single unit running "inference on AI models with up to 200 billion parameters" and fine-tuning "models of up to 70 billion parameters locally," again vendor claims, and it did not repeat the two-unit 405B figure. NVIDIA's product page, checked 28 September 2026, describes a larger configuration: it says ConnectX networking "enables the connection of up to four NVIDIA DGX Spark systems to work with AI models of up to 700 billion parameters," another vendor claim. None of these model-size limits is an independent measurement. Treat each as what NVIDIA said its hardware could do on the date it said it.
NVIDIA's technical documentation lists FP4 compute of "up to 1,000 TOPS" for inference and "up to 1 PFLOP" at FP4 precision with sparsity. Its product page, checked 28 September 2026, footnotes that PFLOP figure as "Theoretical FP4 TOPS using the sparsity feature," confirming it is the sparse number, not the dense one. Treat this vendor claim as a theoretical sparse peak, not measured application throughput.
What it runs, measured
The measurements below were published on 13 and 14 October 2025. Treat them as dated results whose software and setup may differ from yours, not as measurements of today's performance.
LMSYS, publishing through the SGLang team on 13 October 2025, reported gpt-oss-20b results in MXFP4 precision through Ollama on three systems.
| System | Prefill (tokens/sec) | Decode (tokens/sec) |
|---|---|---|
| DGX Spark | 2,053 | 49.7 |
| RTX PRO 6000 Blackwell Workstation Edition | 10,108 | 215 |
| RTX 5090 | 8,519 | 205 |
On this one task, the RTX PRO 6000 decoded about 4.3 times faster than the Spark and the RTX 5090 about 4.1 times faster, both derived from the figures above.
The same LMSYS post shows why batch size matters. Running the larger Llama 3.1 70B in FP8 through SGLang, DGX Spark measured "803 tps prefill / 2.7 tps decode." Running the smaller Llama 3.1 8B in FP8, also through SGLang, tells the other half of the story.
| Batch size | Prefill (tokens/sec) | Decode (tokens/sec) |
|---|---|---|
| 1 | 7,991 | 20.5 |
| 32 | 7,949 | 368 |
Prefill barely moved, but decode rose about 18.0 times from batch 1 to batch 32, derived. LMSYS's own conclusion matches that result: the unified LPDDR5x memory, "offering up to 273 GB/s, shared across both CPU and GPU," is "the key bottleneck in AI inference performance." Across LMSYS, ServeTheHome and StorageReview, the reported bandwidth limitation affects single-stream generation on large dense models more than prefill, batched or mixture-of-experts workloads.
ServeTheHome's 14 October 2025 review, running Ollama out of the box, measured gpt-oss-20b "often over 49 tokens/second," gpt-oss-120b at "14.48 tokens/second," and Qwen3 32B at "9-10 tokens/second." StorageReview, reviewing the device the same day, measured Llama 3.1 8B decode rising from "13.6 tok/s" at BF16 to "23.2" at FP8 to "34.1" at FP4 at concurrency 1, with reported higher-concurrency results of "408.6," "752.8" and "924.1" tok/s respectively. It explicitly gives concurrency 128 for the BF16 result. It put gpt-oss-120b in NVFP4 at "31.4 tok/s at concurrency 1 and 162.7 tok/s at 64 concurrent requests," and measured achieved matrix-multiply throughput of "99.8 TFLOPs" at BF16 and "207.7 TFLOPs" at FP8, both measured rather than theoretical peaks. These results use different setups; do not treat them as a controlled comparison between reviewers.
Reported limitations
Independent reviewers flagged three limits worth knowing before you buy. LMSYS's own read, quoted above, is that 273 GB/s of shared memory bandwidth is a bottleneck for single-stream generation on large dense models. Simon Willison, writing on 15 October 2025, ran into early ARM64 software friction: obtaining PyTorch wheels built for Arm CUDA was difficult, NVIDIA's documentation "changed substantially" within a week of launch, and he judged "It's a bit too early for me to provide a confident recommendation." ServeTheHome, on 14 October 2025, found the dual-port networking undershooting its own spec: combined throughput across both QSFP ports "maxed out at approximately 100Gbps," about half the advertised figure.
Memory bandwidth: the real speed limit
NVIDIA's developer blog, dated 17 November 2023, explains why memory bandwidth can set the pace of text generation. During decoding, it states, "The speed at which the data (weights, keys, values, activations) is transferred to the GPU from memory dominates the latency, not how fast the computation actually happens. In other words, this is a memory-bound operation." The same post describes reading your prompt, the step called prefill, as a highly parallel matrix operation that saturates GPU utilization. This explains the bandwidth comparison below; it is not a runtime prediction for every model or batch size.
DGX Spark's own memory bandwidth is 273 GB/s, per NVIDIA's technical documentation. The site's spec table below shows the corresponding specifications for three GPU families to compare with it.
| Spec | RTX 5090 | RTX PRO 6000 | H100 |
|---|---|---|---|
| VRAM | 32 GB | 96 GB | 80 to 94 GB |
| Memory type | GDDR7 | GDDR7 | HBM3 |
| Memory bandwidth | 1,792 GB/s | 1,792 GB/s | 3,350 GB/s |
| FP4 (dense) | 1,676 TFLOPS | 2,015.2 TFLOPS | Not published |
| FP4 (with sparsity) | 3,352 TFLOPS | 4,030.4 TFLOPS | Not published |
| TDP | 575 W | 600 W | 700 W |
By that table, the RTX 5090 and the RTX PRO 6000 Blackwell each carry 1,792 GB/s of memory bandwidth, about 6.6 times the Spark's 273 GB/s (derived). The H100 family entry gives 3,350 GB/s, about 12.3 times the Spark's figure (derived). Apple's Mac Studio provides another memory-bandwidth comparison. Apple's specification page, checked 28 September 2026, lists the base M5 Max configuration at 460 GB/s, about 1.7 times the Spark's bandwidth (derived), and the same chip configured with a 40-core GPU at 614 GB/s, about 2.2 times (derived). It lists the M5 Ultra at 1.2 TB/s, or 1,200 GB/s, about 4.4 times the Spark's bandwidth (derived). These are specification ratios, not measured application speedups.
The trade-off is now clear: NVIDIA's documented 128 GB of shared system memory exceeds the dedicated GPU memory in each family in the table above, and the base Mac Studio memory capacities Apple lists. Shared system memory is not the same as dedicated GPU memory, so check your model's requirements rather than treating the whole pool as model storage. The Spark's listed bandwidth is the lowest of this comparison group, and LMSYS's gpt-oss-20b results above show slower decode than the two RTX cards on that task. As local LLM hardware, the Spark offers a memory-capacity trade-off, not a general speed advantage. For a longer walkthrough of reading a spec sheet this way, see how to read GPU specs for AI, and check exact model memory needs with the LLM VRAM calculator before you decide.
What that budget buys in rented GPU-hours
At the $4,699 Founders Edition price Tom's Hardware reported on 27 February 2026, this is what the same money buys in rented GPU time. Check NVIDIA's store for today's price before you decide.
| GPU | $/GPU-hr | Provider | GPU-hours for $4,699 | Days at 24 h a day |
|---|---|---|---|---|
| H100 | $2.50 | Hyperstack | 1,879 | 78.3 |
| RTX PRO 6000 Blackwell | $0.59 | RunPod | 7,964 | 331.8 |
| RTX 5090 | $0.53 | Vast.ai | 8,866 | 369.4 |
| RTX 4090 | $0.43 | Vast.ai | 10,927 | 455.3 |
| A100 | $0.68 | LeaderGPU | 6,910 | 287.9 |
Read the table above as an opportunity cost, not a verdict. It converts the purchase budget into GPU-hours; it does not establish how much work your application completes in that time. Before buying, compare your measured workload, schedule and full costs in the rent versus buy GPU calculator. Check the rental's billing and uptime terms rather than assuming every offer works the same way.
Buy the Spark, or rent: five situations
The right answer changes with the job, not just the budget.
| Situation | Buy Spark or rent | Why |
|---|---|---|
| Prototype agents locally | Buy Spark | Choose local development if NVIDIA's documented shared memory suits the model and context you need. Configure the workflow to keep your data on site. |
| Fine-tune a 70B model | Rent first | NVIDIA's 13 October 2025 vendor claim includes 70B fine-tuning, but no published review times that run against an H100 or RTX PRO 6000 Blackwell. Compare your actual run before buying. |
| Serve a model to users | Rent first | Evaluate throughput at your target concurrency and the provider's uptime terms before committing to production. |
| Run overnight batch jobs | Buy Spark if speed can wait | Schedule the job and check the results in the morning, but include electricity in your cost comparison. |
| Data that cannot leave the building | Buy Spark | Keep the workload on your own network and check that its software and integrations meet that requirement. |
The decision rule
Buy a DGX Spark when the bottleneck is local memory capacity and you can live with its bandwidth: local prototyping, models you only need to hold and query occasionally, and workloads you need to keep in the building. Rent first for training, fine-tuning, and serving real traffic, then measure your own workload before buying. Use the published RTX results above as evidence for those specific tests, not as an H100 or RTX 5090 guarantee for every job. The DGX systems explainer is the place to compare other DGX systems, and LLM GPU requirements is worth checking for the model you actually plan to run before you commit to either path.
If you want to try an LLM before buying hardware, rent first to learn what you need. If you already know you need a local AI workstation and the Spark's shared memory and measured performance suit your workload, buy the Spark.
Sources
- NVIDIA DGX Spark hardware documentation, checked 28 September 2026
- NVIDIA DGX Spark product page, checked 28 September 2026
- NVIDIA: Project DIGITS announcement, 6 January 2025
- NVIDIA: DGX Spark launch announcement, 13 October 2025
- Arm Newsroom: Arm-powered NVIDIA DGX Spark AI workstations, 23 March 2026
- Tom's Hardware: DGX Spark price increase, 27 February 2026
- Tom's Hardware: NVIDIA server price warning, 23 August 2026
- TrendForce: HBM3E price hike report, 24 December 2025
- NVIDIA Developer Blog: mastering LLM inference optimization, 17 November 2023
- LMSYS / SGLang team: NVIDIA DGX Spark benchmarks, 13 October 2025
- StorageReview: NVIDIA DGX Spark review, 14 October 2025
- ServeTheHome: NVIDIA DGX Spark review, page 3, 14 October 2025
- ServeTheHome: NVIDIA DGX Spark review, 14 October 2025
- Simon Willison: NVIDIA DGX Spark, great hardware, early days for the ecosystem, 15 October 2025
- Apple Mac Studio technical specifications, checked 28 September 2026