InfiniBand vs RoCE Ethernet: Choose by Workload

InfiniBand and RoCE Ethernet both move GPU traffic between servers. See the generations, provider fabrics, and when a single GPU or node needs neither.

By Faiz Ahmed•
•14 min read

InfiniBand and RoCE Ethernet both carry RDMA traffic between GPU servers, and both run frontier-scale training: Meta reported on 12 March 2024 that it runs one 24,576-GPU cluster on each, with Llama 3 training on the RoCE one. InfiniBand is lossless by design, while RoCE is Ethernet configured not to drop packets, which Meta found took real tuning. For a single GPU or a single eight-GPU server you need neither: vLLM recommends keeping inference inside one GPU or node whenever the model fits.

Everything below is about the network between servers, not inside one. If a single GPU or a single eight-GPU box already does the job, stop reading and go pick a price on the H100 rent page or the GPU hub.

The IBTA says it maintains both specifications and was established in August 1999 by industry leaders; its anniversary page, retrieved 28 September 2026, reports more than 50 members.

Terms, briefly

TermWhat it is
RDMARemote Direct Memory Access: server-to-server data movement directly between application memory, with no CPU involvement on either side (NVIDIA docs, retrieved 28 September 2026).
InfiniBandA switched-fabric interconnect standard defined and maintained by the InfiniBand Trade Association, the IBTA (IBTA, retrieved 28 September 2026).
RoCE v2RDMA over Converged Ethernet, version 2: RDMA carried over UDP and IP, using destination port 4791, so it can cross Layer 3 routers, unlike the Layer-2-only RoCE v1 (NVIDIA docs, retrieved 28 September 2026).
GPUDirect RDMALets a network adapter read and write GPU memory directly, bypassing CPU host memory (NVIDIA docs, retrieved 28 September 2026).
Spectrum-XNVIDIA's name for its AI-optimized Ethernet platform: purpose-built Spectrum switches paired with BlueField or ConnectX SuperNICs (NVIDIA, retrieved 28 September 2026).
Ultra EthernetAn open specification from the Ultra Ethernet Consortium, under the Linux Foundation, for an Ethernet-based high-performance networking stack. The Linux Foundation names AMD, Broadcom, Cisco, Intel, Meta and Microsoft among the founders; specification 1.0 was released on 11 June 2025 (Linux Foundation, 19 July 2023 and 11 June 2025).
EFAAWS's Elastic Fabric Adapter, a network device for EC2 instances that uses the Scalable Reliable Datagram protocol instead of InfiniBand or RoCE directly (AWS docs, retrieved 28 September 2026).

InfiniBand generations and speeds

AscentOptics, an optical-transceiver reseller, gives the following per-port speeds, introduction years and connectors in its compilation retrieved 28 September 2026: 100 Gb/s for EDR (2014), 200 Gb/s for HDR (2017), 400 Gb/s for NDR (2021) and 800 Gb/s for XDR (2024).

GenerationPer-port speedYear introducedConnector
EDR100 Gb/s2014QSFP28
HDR200 Gb/s2017QSFP56
NDR400 Gb/s2021QSFP112 or OSFP
XDR800 Gb/s2024OSFP

NVIDIA's own products track this table closely. Its Quantum-2 switch is the NDR-generation product, and Quantum-X800 is the XDR-generation product, announced on 18 March 2024, about five months after the IBTA announced the XDR specification on 6 October 2023.

The provider documentation below, retrieved 28 September 2026, gives 400 Gb/s per GPU for Azure's ND H100 v5 and ND H200 v5, and a derived average of 400.0 Gb/s for the eight-GPU AWS P5/P5e/P5en and Google Cloud A3 Ultra/A4 configurations. This is a recurring allocation in those examples, not a rule for every Hopper or Blackwell node. Meta's SIGCOMM 2024 paper describes Grand Teton directly: eight GPUs, eight RDMA NICs, and "a 1:1 mapping between GPUs and NICs."

Reading these numbers. Bandwidth shows up in two units that are easy to confuse: Gb/s (gigabits per second) for network ports, and GB/s (gigabytes per second) for NVLink and memory. One GB/s equals 8 Gb/s. A figure can also describe one GPU or an entire node: AWS's P5 documentation, retrieved 28 September 2026, states 3,200 Gbps of EFA networking for an eight-GPU P5/P5e/P5en instance: the derived average is 3,200 / 8 = 400.0 Gbps per GPU. Headline figures also differ on whether both directions are summed. Check the unit, the denominator and the direction before comparing two numbers.

InfiniBand vs RoCE, compared

InfiniBandRoCE v2
Who makes itIBTA defines the standard and lists HPE, IBM, Intel and NVIDIA on its steering committee (retrieved 28 September 2026). NVIDIA's SEC-filed announcement on 11 March 2019 valued the Mellanox acquisition at approximately 6.9 billion dollars in enterprise value; its closing filing is dated 27 April 2020.IBTA also defines RoCE. Dell'Oro's 10 March 2026 report names Celestica, NVIDIA, Arista, Cisco and HPE/Juniper among leading Ethernet-switch vendors for AI back-end networks in 2025.
Flow controlHPCwire describes InfiniBand as intended to be lossless: endpoints receive credits based on receiving-buffer space, and senders transmit only with sufficient credit.NVIDIA recommends Priority Flow Control for RoCE. WWT describes DCQCN as combining PFC with Explicit Congestion Notification to keep queues shallow before PFC must pause traffic.
Tuning it needsAscentOptics describes the Subnet Manager as discovering the fabric and bringing ports to the Active state. NVIDIA's UFM documentation describes its commercial manager as using OpenSM underneath; OpenSM is also an open-source Subnet Manager.Meta's SIGCOMM 2024 paper found DCQCN difficult to tune for training collectives at 200G and 400G. It reported over a year using PFC without other transport-level congestion control for its 400G deployment, while moving congestion management into the collective library.
Documented examplesMeta's 12 March 2024 report: 24,576 GPUs on NVIDIA Quantum-2. Microsoft's ND H100 v5 and ND H200 v5 documentation: Quantum-2 InfiniBand (retrieved 28 September 2026).Meta's 12 March 2024 report: 24,576 GPUs on RoCE, including ongoing Llama 3 training. NVIDIA's 28 October 2024 Colossus announcement: 100,000 Hopper GPUs on Spectrum SN5600 switches with BlueField-3 SuperNICs.
Cost and supplyThe current InfiniBand switches named in this article, Quantum-2 and Quantum-X800, are NVIDIA products.Dell'Oro's 10 March 2026 report attributes part of Ethernet's 2025 demand to buyers wanting more than one supplier while supply was tight.

AWS describes EFA as a network device using Scalable Reliable Datagram transport with OS bypass and congestion control. Treat that as a separate architecture when comparing the documented InfiniBand and RoCE options.

Microsoft Research lists "Congestion Control for Large-Scale RDMA Deployments," the DCQCN paper by Microsoft and Mellanox researchers, at ACM SIGCOMM in August 2015. The paper identifies head-of-line blocking and unfairness with PFC and proposes an end-to-end congestion-control scheme for RoCE v2. Meta's 2024 paper describes DCQCN as established for storage-focused networks but difficult to tune for its AI training collectives; that is one operator's experience, not a general verdict on DCQCN.

Inside the node vs between nodes

NVLink and NVL72 solve a different problem from InfiniBand and RoCE: how GPUs inside one server, or one rack-scale domain, talk to each other. NVIDIA's GB200 NVL72 links 72 Blackwell GPUs into one NVLink domain at 130 TB/s. InfiniBand and RoCE can connect separate domains. NVIDIA's HGX B200 page lists 14.4 TB/s of total NVLink bandwidth and 0.8 TB/s of networking for the eight-GPU system. The two totals are not measured the same way, but the gap is why a job that fits inside one node should stay there. A multi-node cluster can connect separate NVLink domains through InfiniBand, RoCE or another fabric such as AWS's EFA. The NVLink, PCIe and SXM guide covers the inside-the-node side in full, and GB200 and GB300 NVL72 explained covers the rack-scale NVLink domain itself.

What each provider states

The table records what each provider stated on its own pages on 28 September 2026. Use it to frame your questions about the configuration you plan to rent.

ProviderProductFabric statedBandwidth stated
AWSEight-GPU P5 / P5e / P5en configurationsEFA (Elastic Fabric Adapter)3,200 Gbps per instance; derived average 400.0 Gbps per GPU (3,200 / 8)
Microsoft AzureND H100 v5 / ND H200 v5NVIDIA Quantum-2 InfiniBand400 Gbps per GPU, 3.2 Tbps per VM
Google CloudA3 Ultra / A4GPUDirect RDMA over ConnectX-7 NICs3,200 Gbps dedicated to GPU traffic per VM; derived average 400.0 Gbps per GPU (3,200 / 8)
CoreWeaveHGX H100 8-GPU nodesSHARP-enabled NVIDIA Quantum InfiniBand3.2 Tbps per node, bidirectional
Lambda1-Click Clusters, HGX B200 and H100NVIDIA Quantum-2 InfiniBand with SHARPNo Gb/s figure found on the checked page
NebiusGPU clustersNon-blocking NVIDIA Quantum-2 InfiniBandNo Gb/s figure found on the checked pages
CrusoeCloud GPU clustersInfiniBand or high-performance Ethernet, RDMA-backedNo bandwidth figure found on the checked page
Voltage ParkGPU clustersNVIDIA Quantum-2 InfiniBand3,200 Gbps aggregate; per-node/per-GPU basis unspecified
HyperstackH100 SXM / GB200 NVL72, separate pagesGB200 page names Quantum-X800 InfiniBand and Spectrum-X800 EthernetH100 SXM page gives 3.2 Tbps aggregate cluster networking; this is not a GB200 figure
RunPodInstant ClustersInfiniBand or RoCE v2, by configuration1,600 to 3,200 Gbps east-west, headlined as 3,200 Gbps; per-node/per-GPU basis unspecified
DenvrGPU clustersNon-blocking InfiniBand3,200 Gbps; per-node/per-GPU basis unspecified
Latitude.sh8x HGX B300 bare metalRoCE, dual plane800 Gbps for this SKU, plane/direction basis unspecified; no cluster fabric found in the checked H100 and RTX PRO 6000 listings
Hot AisleAMD GPU cloudRoCEv28x400G; the page does not explicitly map this to GPUs
VultrCloud GPUNot stated on the pages checkedNot stated on the pages checked

A blank means no figure appeared on the provider's public pages that day, so ask. Where the denominator or direction is unspecified, ask the provider before comparing the figure with a per-GPU rate.

When the fabric actually matters

Whether any of this matters comes down to the job, not the GPU count alone.

  • Single-node inference. vLLM says distributed inference is probably unnecessary if the model fits on one GPU, and recommends tensor parallelism when it fits in one node.
  • Fine-tuning across several nodes. Either fabric works. Meta reported on 12 March 2024 that both of its 24,576-GPU clusters ran large GenAI workloads without network bottlenecks. Ask the provider for results on a job like yours rather than choosing by fabric name.
  • Pre-training at scale, tens of thousands of GPUs. Meta's March 2024 report describes its paired 24,576-GPU clusters. NVIDIA's 28 October 2024 Colossus announcement describes 100,000 Hopper GPUs for Grok training, connected through Spectrum SN5600 switches and BlueField-3 SuperNICs. NVIDIA says the buildout took 122 days and training began 19 days after the first rack arrived. These are vendor-reported deployment details. Expect the provider to name a specific fabric and generation, not just high-speed networking.
  • Multi-node inference for very large mixture-of-experts models. vLLM documents expert parallelism across GPUs and multi-node communication through DeepEP. DeepEP, used by vLLM and SGLang, describes NVLink within a node and RDMA between nodes, and says it is tested with InfiniBand. Its project benchmark uses different GPU generations for the NVLink and RDMA results, so it does not isolate the effect of the fabric. For this deployment, ask how the provider supports the communication path your serving stack uses.

For a job that stays within one node, start with model fit and price. For multi-node work, use the cluster page and ask for the fabric details alongside the quote.

The GPUs multi-node clusters are built from

The provider descriptions above include H100, H200 and B200 systems. The live table below compares the SXM variants; use the cluster page when your job needs more than one node, and ask for a quote tied to the requested topology.

GPUCheapest $/GPU-hrProviderProviders in stock
H100 SXM5$2.90Ori4
H200 SXM$3.50Ori4
B200 SXM$7.20VERDA1
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

Before you rent a multi-node cluster

  • Which fabric, and which generation. "InfiniBand" or "high-speed networking" alone is not an answer. Ask whether it is NDR, XDR, Spectrum-X800 or something newer.
  • Bandwidth per GPU, not just per node. AWS documents 3,200 Gbps for the eight-GPU P5/P5e/P5en configurations. Confirm the denominator before dividing; an aggregate does not by itself establish a dedicated connection per GPU.
  • Whether GPUDirect RDMA is enabled for the specific instance size you would rent, not just the GPU generation. AWS's own EFA documentation states this varies by instance size within the same family.
  • How congestion is controlled, and whether it has been tuned for collective training traffic specifically. Meta's 2024 paper distinguishes transport-level PFC from the congestion management it moved into the collective library; ask how those layers work together in your configuration.
  • Whether the fabric is non-blocking across the whole cluster or segmented into zones that need a higher switch tier to talk to each other.
  • How the provider handles Subnet Manager failure and recovery. AscentOptics explains that a Subnet Manager brings InfiniBand ports to the Active state; ask what failover protects that process.

The decision rule: keep inference inside one GPU or node whenever it fits. For multi-node training, either fabric can do the job, so choose the provider on price and availability, and insist on a named fabric generation, a per-GPU bandwidth figure and results on a workload like yours before you commit.

Sources

Frequently asked questions

What is the difference between InfiniBand and RoCE?▾

InfiniBand is a separate network standard, maintained by the InfiniBand Trade Association and lossless by design. RoCE carries the same RDMA traffic over Ethernet, which has to be configured not to drop packets, usually with Priority Flow Control. Meta reported in 2024 that it runs one 24,576-GPU cluster on each.

Do I need InfiniBand or RoCE for a single GPU?▾

No. The fabric only matters between servers. vLLM's documentation says distributed inference is probably unnecessary when a model fits on one GPU, and recommends tensor parallelism inside one node when it fits there.

What is GPUDirect RDMA?▾

NVIDIA defines GPUDirect RDMA as direct access to GPU memory by network devices, bypassing CPU host memory. Support varies by instance: AWS lists it per instance size, so check the exact instance you rent.

What is Spectrum-X?▾

NVIDIA describes Spectrum-X as its Ethernet platform for AI clusters, pairing Spectrum switches with BlueField or ConnectX SuperNICs. Its 28 October 2024 Colossus announcement described 100,000 Hopper GPUs connected through Spectrum SN5600 switches and BlueField-3 SuperNICs.

Does NVLink replace InfiniBand?▾

No. NVLink connects GPUs inside a server, or inside one rack-scale domain such as the 72-GPU GB200 NVL72. InfiniBand or RoCE connects those servers or racks to each other.

What is Ultra Ethernet?▾

The Linux Foundation names AMD, Broadcom, Cisco, Intel, Meta and Microsoft among the Ultra Ethernet Consortium founders. Its 11 June 2025 release announced specification 1.0, emphasizing open standards and interoperability for Ethernet and IP.

Related Posts