NVIDIA Vera CPU Replaces Grace and Goes Standalone

Compare Vera with Grace, including conflicting memory figures, standalone server announcements, reported deliveries and Grace-based GPU rental options.

By Faiz Ahmed•
•10 min read

The NVIDIA Vera CPU has 88 NVIDIA-designed Olympus cores and 176 threads, according to NVIDIA's 5 January 2026 technical article, updated 16 March. It replaces Grace as the CPU in the Vera Rubin platform described by NVIDIA in that article. NVIDIA announced standalone Vera servers and a CPU-only rack on 16 March 2026, then named standalone server suppliers on 31 May.

That standalone option is the practical news: a Vera evaluation for CPU work need not start with a Rubin GPU purchase. But separate three decisions before choosing a system: the CPU design, the server configuration and the access you can actually arrange. A customer delivery announcement answers a different question from a public rental listing.

Vera vs Grace: compare one CPU at a time

NVIDIA's January technical comparison supplies the baseline below. Its March update matters: these are the article's published figures, not a frozen record of the January launch. Architecture details come from NVIDIA's separate Grace and Vera technical articles. Where another publisher disagrees, both figures stay visible.

SpecificationGraceVera
Cores and design72 Arm Neoverse V2 cores, NVIDIA, 5 January 2026, updated 16 March88 NVIDIA Olympus cores, same NVIDIA article
Threads72, NVIDIA, 5 January 2026, updated 16 March176 through Spatial Multithreading, same NVIDIA article
Arm architectureArmv9-A, NVIDIA, 20 January 2023Compatible with Arm v9.2, NVIDIA, 16 March 2026
Memory and published maximum capacity per CPUUp to 480 GB LPDDR5X, NVIDIA, 5 January 2026, updated 16 MarchUp to 1.5 TB LPDDR5X, same NVIDIA article; Micron says its SOCAMM2 portfolio enables up to 2 TB per CPU, 16 March 2026
Memory bandwidthUp to 512 GB/s, NVIDIA, 5 January 2026, updated 16 MarchUp to 1.2 TB/s, same NVIDIA article
NVLink-C2C bandwidth900 GB/s, NVIDIA, 5 January 2026, updated 16 March1.8 TB/s, same NVIDIA article; The Register identifies this as bidirectional, 1 August 2026
On-chip coherency fabricNot compared3.4 TB/s SCF bisection bandwidth, NVIDIA technical article, 16 March 2026; 3.6 TB/s on-chip fabric, NVIDIA COMPUTEX recap, 31 May 2026
ProcessNot comparedANALYST ESTIMATE: apparently TSMC 3nm for the compute die, reported with qualification by The Register, 1 August 2026; not a foundry confirmation
PowerNVIDIA gives a 500 W TDP including memory for the two-CPU Grace CPU Superchip, 20 January 2023; not a single-CPU ratingConfigurable 250 W to 450 W TDP, NVIDIA, 31 May 2026

The power row deliberately stops short of an efficiency comparison. NVIDIA's 20 January 2023 Grace article describes a two-CPU Superchip, while its 31 May 2026 Vera article gives Vera's configurable TDP. Dividing the Grace rating in half would not establish a published single-CPU rating. Nor would comparing those limits tell you the energy required to finish your job.

The same discipline applies to memory. Keep per-CPU capacity separate from a whole server's total. Ask for the installed memory configuration when requesting a quote, and use that figure in your job plan. A maximum supported configuration is not a promise about the machine you will receive.

NVIDIA Olympus changes the core design

NVIDIA's 16 March 2026 technical article describes Olympus as its own Arm-compatible design. That is one clear Grace vs Vera distinction: NVIDIA's 20 January 2023 Grace article identifies Arm Neoverse V2 cores, whereas the Vera article identifies NVIDIA-designed cores compatible with Arm v9.2.

NVIDIA's March Vera article also says Spatial Multithreading lets users choose performance per thread versus thread count at runtime. Treat thread count as a configuration detail to test against your workload. Do not turn the table's thread totals into an assumed application speedup. Request the intended threading configuration before comparing results from different machines.

NVIDIA's March article places Vera's cores on a monolithic compute die, with memory and I/O on adjacent dielets. It also describes detachable, upgradable SOCAMM memory modules. Those details matter when specifying a server: ask which memory configuration the supplier qualifies and what an upgrade requires. The package description alone does not answer either purchasing question.

Two conflicts to resolve before ordering

NVIDIA's January comparison lists up to 1.5 TB of LPDDR5X per Vera CPU. Micron's 16 March 2026 announcement says its SOCAMM2 portfolio enables up to 2 TB per Vera CPU. These statements describe different published limits. Do not replace NVIDIA's configuration with Micron's larger figure without confirmation from the server supplier.

There is also a disagreement inside NVIDIA's own publications. Its 16 March technical article specifies 3.4 TB/s of Scalable Coherency Fabric bisection bandwidth. Its 31 May COMPUTEX recap prints 3.6 TB/s for Vera's on-chip fabric. The precise metric accompanies the former; the latter does not resolve the discrepancy. Neither should be substituted for the separate NVLink-C2C row.

For your specification, write down the memory capacity, memory bandwidth and CPU link separately. Require the supplier to identify the applicable configuration. That avoids making a buying decision around an attractive number whose scope differs from the system being quoted.

Vera inside Rubin is one configuration

NVIDIA's Vera Rubin datasheet, as of 1 October 2026, pairs one Vera CPU with two Rubin GPUs; NVIDIA's DGX page on the same date lists 36 Vera CPUs and 72 Rubin GPUs in Vera Rubin NVL72. NVIDIA's updated 13 October 2025 blog records the rename from Vera Rubin NVL144 to NVL72, and SemiAnalysis explained on 25 February 2026 that the name now counts GPU packages instead of compute dies. See the Vera Rubin platform guide for the rack design and the NVIDIA GPU roadmap for the wider generation sequence.

Standalone Vera has its own delivery story

NVIDIA's 16 March 2026 launch introduced a CPU-only rack containing 256 liquid-cooled Vera CPUs, built on its MGX modular reference architecture. NVIDIA's technical article that day also listed standalone single- and dual-socket Vera servers. These are distinct CPU system configurations, so a Vera evaluation need not start with the GPU rack.

The March launch named Alibaba Cloud, ByteDance, Meta, OCI, CoreWeave, Lambda, Nebius and Nscale as deployment collaborators. NVIDIA's 31 May announcement then explicitly named Dell, HPE, Lenovo and Supermicro as suppliers of standalone Vera CPU servers. It described Vera as being in full production and named Anthropic, OpenAI and SpaceXAI among labs planning adoption.

Read those names with their dates and roles attached. A deployment collaborator is not automatically a public rental provider. A supplier announcement is a reason to request a configuration and delivery commitment, not a substitute for either. NVIDIA's 31 May announcement establishes the standalone server route; it does not establish a boxed retail CPU sales channel.

NVIDIA's delivery article, first published 18 May and updated 27 August 2026, reports initial Vera CPU system deliveries to Anthropic, OpenAI, SpaceXAI and OCI. The August update also reports delivery of AWS's first Vera CPU server. These are reported deliveries by NVIDIA, not evidence that an arbitrary customer can provision the same system immediately. The same article says OCI plans to deploy hundreds of thousands of Vera CPUs beginning in 2026. On 22 June 2026, NVIDIA announced that Los Alamos National Laboratory's Veritas system, to be built by HPE, includes standalone Vera CPU partitions.

For a purchase, use the named server suppliers as the starting point for a quote. For a rental, require an actual offer with the CPU allocation, memory and access terms specified. Do not use a lab's adoption plan or another customer's reported delivery as your project's start date.

When the CPU deserves attention in an AI job

NVIDIA's 31 May 2026 standalone announcement positions Vera for agentic AI, reinforcement learning and data processing. The CPU therefore deserves attention when your job includes those workloads. The useful question for your own evaluation is how much of the end-to-end job they consume.

The intended workload range extends beyond that positioning. In Tom's Hardware's 23 March 2026 GTC press Q&A transcript, NVIDIA's Ian Buck described the dual-socket reference system as supporting PyTorch, compilation, SQL and HPC workloads. That is a statement about intended workload support, not a measured ranking against another CPU.

Evaluate the stage you need to improve. Record its elapsed time, memory and concurrency on your current machines, then run the same job, with the same completion criteria, on the Vera configuration a supplier intends to deliver, including its memory and threading mode. A faster stage is worth paying for only if it shortens the whole job enough to cover the system's cost.

Read Grace-based rental listings separately

For rental shopping, the live table below shows the Grace-based GH200 and GB300 first, with H200 and B200 listings for comparison. ASUS's GH200 server page, as of 28 September 2026, identifies its system as Grace Hopper. NVIDIA's GB300 page, as of 28 September 2026, describes a rack combining Grace CPUs with Blackwell Ultra GPUs. Neither name means that the CPU is Vera.

GPUCheapest $/GPU-hrProviderProviders in stock
GH200 Grace Hoppernone in stock
GB300none in stock
H200$3.43QuantaCloud7
B200$7.20VERDA1
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Read each row as the live per-GPU rental comparison, then inspect the offer's CPU and host-memory allocation before treating it as a quote for your complete job.

Open the GH200 rental listings and GB300 rental listings to compare the offers. GB200 also belongs in the Grace discussion: NVIDIA's GB200 page, as of 28 September 2026, describes Grace CPUs paired with Blackwell GPUs. The GB200 and GB300 system guide explains those configurations.

When assessing a listing, request the CPU resources your process can use, the host memory assigned to it and the relevant interconnect details. Use the NVLink, PCIe and SXM guide to interpret that terminology. Keep your comparison tied to the offered instance rather than assuming every listing exposes an entire system's resources.

Choose Vera when you have identified CPU work worth accelerating and can obtain a specified server configuration to evaluate. If you need GPU time now, compare the live Grace-based offers and choose against your job's measured requirements. Commit to Vera only after its CPU benefit and access terms solve a concrete problem for that job.

Sources

Frequently asked questions

What is the NVIDIA Vera CPU?▾

NVIDIA's January 2026 technical article describes Vera as an 88-core CPU with NVIDIA-designed Olympus cores and 176 threads. It succeeds Grace in the Vera Rubin platform.

What does NVIDIA Olympus mean?▾

Olympus is NVIDIA's CPU core design inside Vera. NVIDIA's 16 March 2026 technical article specifies compatibility with Arm v9.2.

Can Vera be bought without Rubin GPUs?▾

NVIDIA announced standalone Vera servers and a CPU-only rack on 16 March 2026. Its 31 May announcement named Dell, HPE, Lenovo and Supermicro as standalone server suppliers; that does not establish boxed retail CPU sales.

Does Vera support 1.5 TB or 2 TB of memory?▾

NVIDIA's January 2026 comparison lists up to 1.5 TB of LPDDR5X per CPU. Micron's 16 March 2026 announcement says its SOCAMM2 portfolio enables up to 2 TB per Vera CPU; verify the installed and qualified capacity of the specific server.

Can I rent a standalone Vera CPU?▾

Not from the GPU rental listings this site tracks, as of 1 October 2026. NVIDIA has announced standalone Vera servers and reported deliveries to AI labs and clouds; for GPU rentals today, compare the Grace-based GH200 and GB300 listings.

Related Posts