NVIDIA announced Rubin CPX on 9 September 2025 as a context, or prefill, GPU, then pulled it from its 2026 plans: Ian Buck said "we've pulled CPX" at the GTC press Q&A, whose transcript Tom's Hardware published on 23 March 2026. NVIDIA Groq 3 LPX, built on language processing unit technology licensed from Groq, joined the Vera Rubin platform as an additional inference accelerator; NVIDIA announced full production on 24 August 2026. As of 1 October 2026, neither CPX nor LPX has a confirmed public rental or callable endpoint: Nebius and Groq announced LPX deployment plans on 24 August, but neither announcement establishes public access.
For an inference team, the useful change is the serving design. Separate the work that reads a prompt from the work that writes an answer. Test that separation in software before making a hardware commitment.
Disaggregated inference starts with two different waits
Prefill reads the input. NVIDIA's 9 September 2025 technical explanation describes it as compute-intensive context processing that produces the first output token. Think of the pause after you submit a long document. Its proposed disaggregated design assigns this work its own resources, with low-latency transfer of the KV cache, the attention state passed to the next phase.
Decode produces subsequent tokens. The same NVIDIA explanation calls this phase bandwidth-sensitive. Think of the pace at which the answer appears after it starts. NVIDIA separates the phases so compute and memory resources can be tuned independently. These are two parts of serving one request, not a division between training and serving; the training versus inference guide covers that distinction.
Rubin CPX: the announced design and the change of plan
NVIDIA's 9 September 2025 announcement positioned CPX for million-token coding and generative-video context processing. It described a monolithic Rubin-architecture die with 128GB of GDDR7 per GPU. That is the original NVIDIA Rubin CPX design, not a specification for a shipping product. The HBM versus GDDR comparison explains the memory distinction.
VENDOR CLAIM: NVIDIA's same announcement advertised up to 30 petaflops at NVFP4 precision per CPX GPU. The announcement does not label that figure dense or sparse. Do not divide that headline by a Blackwell specification and treat the result as application speed.
The original integrated rack was named Vera Rubin NVL144 CPX. NVIDIA's 9 September release specified 100TB of fast memory. VENDOR CLAIM: it advertised 8 exaflops of AI compute, 1.7 petabytes per second of memory bandwidth and 7.5 times the AI performance of GB300 NVL72. The release does not label the rack compute figure dense or sparse. Those are announcement claims, not measurements of your model.
Rack naming needs care. NVIDIA's editorial note on its 13 October 2025 blog records a branding change from Vera Rubin NVL144 to NVL72. SemiAnalysis explained on 25 February 2026 that the old name counted 144 compute dies in 72 two-die packages; the new name counts packages. Keep the CPX rack's historical name when discussing its announcement. Do not read its number as a current package count.
NVIDIA originally targeted CPX availability for the end of 2026. That plan changed. In the transcript published on 23 March 2026, Buck attributed the CPX decision to prioritizing LPU decode delivery with Vera Rubin during 2026. Asked about CPX shipment that year, he said: "I'm never going to say no to how fast we can innovate." That leaves room for later work, not a delivery commitment.
ANALYST ESTIMATE: Investing.com reported on 31 August 2026 that Ming-Chi Kuo's industry checks indicated a CPX restart with production in the first quarter of 2027. NVIDIA has not confirmed that revival as of 1 October 2026. Do not reserve a launch dependency against it. The NVIDIA GPU roadmap keeps the wider timeline separate from this serving decision.
LPX has more than one documented decode role
NVIDIA announced NVIDIA Groq 3 LPX on 16 March 2026 as a fully liquid-cooled MGX inference system, targeting availability in the second half of 2026. Its technical article that day described 256 interconnected LPU accelerators and 128GB of total on-chip SRAM per rack. NVIDIA's LPX page, as of 1 October 2026, also lists 12TB of DDR5 per rack. Calling the system SRAM-only would be wrong.
NVIDIA's 16 March architecture article describes an attention and feed-forward split. Rubin GPUs retain prefill and decode attention. LPX runs the latency-sensitive feed-forward network, or FFN, and mixture-of-experts, or MoE, work within decode. The GPU and LPU exchange intermediate activations for each token. In this design, moving work to LPX does not remove Rubin from the decode loop.
NVIDIA's 24 August technical article describes another arrangement. Vera Rubin performs prefill, then transfers the KV cache to LPX, which performs the entire decode. The distinction matters when evaluating an implementation: ask whether it transfers state between phases or exchanges activations during every token. Saying LPX handles decode does not tell you which architecture is being offered.
NVIDIA's 16 March POD description gives LPX direct chip-to-chip links through a copper rack spine, extendable across LPX racks. It describes rack-to-rack connectivity through Spectrum-X Ethernet or Quantum-X800 InfiniBand networking racks. LPX operates beside the GPU racks; the Vera Rubin platform guide covers the rest of that system.
NVIDIA's 24 August 2026 production announcement still described Nebius Token Factory delivery as planned. Nebius's announcement that day said its initial service would cover a subset of models. Groq's announcement that day named Dell as its deployment partner and described future access through its platform, APIs and enterprise services. A production milestone is not an endpoint you can put into a deployment configuration.
The NVIDIA Groq deal licensed technology and announced staff moves
Groq announced the non-exclusive inference-technology licence on 24 December 2025. It said founder Jonathan Ross, president Sunny Madra and other team members would join NVIDIA. Groq also said it would remain independent and GroqCloud would continue without interruption. Groq's later statement on 12 August 2026 identifies Adam Winter as CEO.
The technology connection is explicit. NVIDIA's annual report filed on 25 February 2026 identifies the licensed subject as "language processing unit technology." Buck also discussed licensing the IP when explaining Groq 3 integration in the March Q&A. This was more than permission to use a name.
Reported deal value: CNBC reported $20 billion in cash, relayed by Reuters on 24 December 2025; Reuters noted that neither company commented on that report. NVIDIA's annual report filed on 25 February 2026 instead discloses $13.0 billion paid at closing plus $4 billion payable within one year, including imputed interest. It says NVIDIA bought no Groq customer contracts, existing products or equity interests. These disclosures do not support calling the NVIDIA Groq deal an acquisition of Groq. The filing is the primary source for what NVIDIA paid.
Reuters on 9 September 2026 relayed a New York Times report of a DOJ inquiry into whether the deal was structured to avoid antitrust scrutiny. Reuters said it could not immediately verify the report. This is a reported inquiry, not a finding of wrongdoing.
For an existing Groq API user, keep the service and hardware roadmap separate. Groq's December continuity statement does not establish LPX backing for a particular request. The accelerator and API guide covers Groq's API alongside other alternatives.
Four choices, four different commitments
The table separates hardware status from the job you assign it. Rubin's HBM4 is specified in NVIDIA's 21 July 2026 architecture article. B200's HBM3e entry comes from this site's NVIDIA-sourced specification table. The GPU-only split is a deployment option documented by Dynamo, not a separate accelerator product.
| Hardware | Job in the pipeline | Memory type | Status and when |
|---|---|---|---|
| Original Rubin CPX | Specialized context/prefill GPU | GDDR7 | NVIDIA announced 9 September 2025; Buck's withdrawal explanation published 23 March 2026 |
| NVIDIA Groq 3 LPX | Decode FFN/MoE beside GPU attention, or whole decode after GPU prefill | On-chip SRAM plus rack DDR5 | NVIDIA announced full production 24 August 2026; public endpoint unconfirmed as of 1 October 2026 |
| Standard Rubin GPU | Prefill and decode attention in NVIDIA's split; prefill in its whole-decode offload design | HBM4 | CoreWeave announced cloud availability 30 September 2026; NVIDIA called access early access that day |
| Blackwell B200 today | GPU serving; separate GPU pools when configured through software | HBM3e | Current rental listings and stock are shown live below |
CoreWeave's 30 September 2026 Rubin announcement and NVIDIA's same-day early-access description concern Rubin GPU systems. They do not establish CPX or LPX access. Ask for the exact accelerator behind a proposed service, not just the platform name.
Use NVIDIA Dynamo, SGLang or vLLM to test the split
NVIDIA's Dynamo documentation, as of 1 October 2026, describes orchestration for disaggregated serving, KV-aware routing, cache management and autoscaling. It names SGLang, vLLM and TensorRT-LLM as supported engines. Dynamo's README, as of 28 September 2026, describes independently scalable prefill and decode GPU pools. It also says an engine alone is probably sufficient for one model on one GPU.
SGLang's documentation, as of 1 October 2026, describes separate prefill and decode servers with a router, using Mooncake or NIXL transfer engines. vLLM's guide, as of the same date, documents separate instances and KV-transfer connectors, including NIXL and Mooncake. These are software paths to evaluate on rented GPU infrastructure. Read the vLLM, SGLang and TensorRT-LLM comparison when choosing the underlying engine.
Keep vLLM's warning attached to the feature: its guide, as of 1 October 2026, calls disaggregated prefilling experimental and explicitly says it does not improve throughput. Its stated purpose is independent tuning of time to first token and inter-token latency. Do not make a capacity forecast by borrowing an LPX performance claim.
The live table below shows current H100, H200 and B200 rental offers for a GPU-only trial.
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| H100 | $2.50 | Hyperstack | 9 |
| H200 | $3.43 | QuantaCloud | 6 |
| B200 | $6.79 | RunPod | 3 |
Start with a baseline on your chosen model and engine. Record prompt lengths, output lengths, concurrency, time to first token and inter-token latency. Then compare a separated deployment at the same resource budget. Include transfer time and idle workers in the evaluation. Keep the split only if it improves the latency target you actually need while preserving acceptable throughput and cost.
For the next 12 months, build around a GPU deployment you can validate now. Add software disaggregation when your trial justifies it. Evaluate LPX when a provider gives you a confirmed model endpoint or hardware allocation and you can repeat that trial. Keep CPX out of committed capacity plans until NVIDIA confirms a product and delivery path.
Sources
- NVIDIA CPX announcement, 9 September 2025.
- NVIDIA CPX technical explanation, Japanese-language page, 9 September 2025.
- Tom's Hardware: GTC press Q&A transcript, 23 March 2026.
- NVIDIA rack branding note, 13 October 2025.
- SemiAnalysis: Vera Rubin design and naming, 25 February 2026.
- Investing.com: Kuo's reported CPX revival, 31 August 2026, unconfirmed analyst report.
- Hardware Busters: unconfirmed CPX revival, 1 September 2026.
- NVIDIA Vera Rubin and LPX announcement, 16 March 2026.
- NVIDIA LPX architecture, 16 March 2026.
- NVIDIA LPX product page, as of 1 October 2026.
- NVIDIA Vera Rubin POD networking, 16 March 2026.
- NVIDIA LPX decode designs, 24 August 2026.
- NVIDIA LPX full-production announcement, 24 August 2026.
- Nebius LPX deployment plans, 24 August 2026.
- Groq LPX deployment plans, 24 August 2026.
- Groq licensing agreement, 24 December 2025.
- Groq NVIDIA Cloud Partner announcement, 12 August 2026.
- Reuters via Investing.com: CNBC's reported deal value, 24 December 2025.
- NVIDIA fiscal-2026 annual report, filed 25 February 2026.
- Reuters via Investing.com: reported DOJ inquiry, 9 September 2026.
- NVIDIA Rubin architecture, 21 July 2026.
- CoreWeave Rubin availability announcement, 30 September 2026.
- NVIDIA CoreWeave early-access description, 30 September 2026.
- NVIDIA Dynamo documentation, as of 1 October 2026.
- Dynamo README, as of 28 September 2026.
- SGLang disaggregation guide, as of 1 October 2026.
- vLLM disaggregated-prefill guide, as of 1 October 2026.