NVIDIA GPU Generations Explained: Decode Any GPU Name

A is Ampere, L is Ada, H is Hopper, B is Blackwell. A decoder table, a generation timeline and a four-step method for placing any NVIDIA GPU on a rental site.

By Faiz Ahmed
13 min read

An NVIDIA data center GPU name is a letter for the architecture, a number for the tier inside it, and a suffix for the form factor. A is Ampere, L is Ada Lovelace, H is Hopper and B is Blackwell. GB in front means Grace CPUs are packaged with Blackwell GPUs. Once you know the generation and the tier, you know which number formats the card accelerates, whether it has NVLink and MIG, and roughly where it will sit in a price table.

The decoder table

Letter or suffixMeaningExample
PPascal data center partP100
VVoltaV100
TTuringT4
AAmpereA100, A40, A10
LAda Lovelace data center partL4, L40S
HHopperH100, H200
BBlackwell, and Blackwell Ultra for B300B200, B300
GBGrace CPUs packaged with Blackwell GPUs (a superchip)GB200, GB300
GH200Hopper-generation part; our table lists its interconnect as NVLink-C2CGH200
GeForce RTX 20, 30, 40, 50 seriesConsumer cards: Turing, Ampere, Ada Lovelace, Blackwell in that orderRTX 3090, RTX 4090, RTX 5090
GTXOlder consumer cards; the Pascal ones we track have no Tensor coresGTX 1080
RTX A plus a numberAmpere workstation cardRTX A6000
RTX, a number, then "Ada"Ada Lovelace workstation cardRTX 6000 Ada
RTX PROBlackwell workstation and server cardRTX PRO 6000 Blackwell
QuadroWorkstation brand; in our table it runs from Maxwell to TuringQuadro RTX 8000
SXMSocketed module on an HGX baseboard, full NVLinkH100 SXM5
PCIeStandard add-in cardA100 PCIe
NVL (on a card)PCIe card joined to neighbours by NVLink bridgesH100 NVL, H200 NVL
NVL8, NVL72Number of GPUs in one NVLink domainGB200 NVL72
HGX, DGXHGX is the eight-GPU baseboard sold to server makers; DGX is NVIDIA's own complete system built on itHGX B200, DGX B200

Two notes on SXM. NVIDIA does not publish what the letters stand for, so we do not expand it. The module is numbered by generation: Wikipedia's SXM article (read 21 September 2026) pairs SXM2 with P100 and V100, SXM4 with A100, SXM5 with H100 and H200, and SXM6 with B200. NVIDIA's own pages say only "Blackwell SXM"; the SXM6 label comes from server makers such as Dell and Lenovo.

The number after the letter ranks cards inside a generation only loosely. An L4 is a 72 W card with 24 GB and an L40S is a 350 W card with 48 GB, but an H200 is not twice an H100. It is the same GPU with more and faster memory, which we cover in how to read GPU specs for AI.

NVIDIA data center GPU list by generation

This list is built from the launch years and architectures in our datasheet-verified spec table. It covers the GPUs this site tracks, not everything NVIDIA has shipped. The full GPU spec chart has every row.

ArchitectureLaunch years in our tableGPUs we track
Pascal2016 to 2017P100, GTX 1070, GTX 1080, Quadro P4000, P5000 and P6000, TITAN Xp
Volta2017V100, TITAN V
Turing2018 to 2019T4, Quadro RTX 4000, 5000, 6000 and 8000, RTX 2060, 2070 and 2080
Ampere2020 to 2021A100, A40, A30, A10, A16, RTX A2000 to A6000, RTX 3060 to 3090
Ada Lovelace2022 to 2024L4, L40, L40S, RTX 2000 Ada to RTX 6000 Ada, RTX 4060 to 4090
Hopper2022 to 2024H100 (2022), GH200 (2023), H200 (2024)
Blackwell2024 to 2025B200 (2024), RTX 5060 to 5090, RTX PRO 6000 Blackwell (2025)
Blackwell Ultra2025B300, GB300

What each generation rents for now

GPUCheapest $/GPU-hrProviderProviders in stock
V100$0.83Ori2
A100$0.68LeaderGPU10
L40S$0.80Vast.ai7
H100$2.50Hyperstack10
H200$3.43QuantaCloud7
B200$3.75Packet.ai1
B300$7.89RunPod1
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

What changed each generation, for a renter

SpecV100A100H100B200
Launch year2017202020222024
ArchitectureVoltaAmpereHopperBlackwell
VRAM16 to 32 GB40 to 80 GB80 to 94 GB180 to 192 GB
Memory bandwidth900 GB/s2,039 GB/s3,350 GB/s8,000 GB/s
FP16 (dense)125 TFLOPS312 TFLOPS989 TFLOPS2,250 TFLOPS
FP16 (with sparsity)Not published624 TFLOPS1,979 TFLOPS4,500 TFLOPS
FP8 (dense)Not publishedNot published1,979 TFLOPS4,500 TFLOPS
FP8 (with sparsity)Not publishedNot published3,958 TFLOPS9,000 TFLOPS
FP4 (dense)Not publishedNot publishedNot published9,000 TFLOPS
FP4 (with sparsity)Not publishedNot publishedNot published18,000 TFLOPS
Figures from the vendor datasheets: V100, A100, H100, B200, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

Pascal (2016). No Tensor cores. The P100 already has the shape of the later data center parts: HBM2 memory, NVLink and an SXM module. Treat Pascal as legacy.

Volta (2017). The first Tensor cores: the V100 has 640 of them beside 5,120 CUDA cores, and with them came FP16 mixed-precision training. It has 16 or 32 GB of HBM2 and second-generation NVLink, which NVIDIA's V100 datasheet lists at 300 GB/s on the SXM2 module. Software support is thinning. NVIDIA's TensorRT support matrix (8 September 2026) no longer lists Volta at all.

Turing (2018). Adds INT8 and INT4 to the Tensor cores, according to SemiAnalysis's history of the design (23 June 2025). The T4, a 70 W card with 16 GB of GDDR6, is the Turing data center part we track. The TensorRT matrix gives it FP32, FP16 and INT8 only, and marks BF16 as not available.

Ampere (2020). Three things arrive at once: BF16 and TF32, MIG partitioning, and third-generation NVLink at 600 GB/s per GPU on the A100 datasheet. BF16 is the one that matters most for training, for reasons covered in FP8 vs FP16 vs BF16. Ampere has no FP8. It is the oldest generation we would choose for new training work. The A100 vs H100 comparison shows what the step up costs today.

Ada Lovelace (2022) and Hopper (2022). These two are the same age and both add FP8; NVIDIA's Transformer Engine needs compute capability 8.9 or higher, which means Ada, Hopper or Blackwell. They split by memory and interconnect. Hopper is the HBM part, with 80 GB of HBM3 on the H100, fourth-generation NVLink at 900 GB/s per GPU and PCIe Gen5, followed by 141 GB of HBM3e on the H200 two years later. Ada is the PCIe part: GDDR6 memory on a PCIe Gen4 card, and on the L40S a block of RT cores for graphics beside the Tensor cores. On the H100 and the L40S, the TensorRT matrix marks FP4 as "hardware emulation mode", so the hardware does not accelerate it.

Blackwell (2024) and Blackwell Ultra (2025). Adds FP4. SemiAnalysis lists NVIDIA's own NVFP4 format and the MXFP8, MXFP6 and MXFP4 microscaling formats as new in this generation. It also brings fifth-generation NVLink at 1,800 GB/s per GPU on NVIDIA's NVLink page (read 21 September 2026). The B200 carries 180 to 192 GB of HBM3e at 8,000 GB/s and a 1,000 W TDP. FP4 also reaches the consumer and workstation cards: NVIDIA lists it for both the RTX 5090 and the RTX PRO 6000 Blackwell, which use GDDR7. B300 is the Blackwell Ultra part, at 1,400 W. Its memory is quoted two ways: on 21 September 2026 Dell listed 270 GB per GPU and Latitude.sh listed 288 GB. No source we found explains the gap, so check the figure on the listing you rent. The B200 vs H200 comparison tracks the price gap between the two newest generations.

All the NVLink figures above come from NVIDIA's datasheets and NVLink page as read on 21 September 2026, and they are totals for both directions added together. Per direction is half, which is why Lenovo's B200 page (updated 13 January 2026) lists 900 GB/s.

Superchips: GH200, GB200 and GB300

The GH200 sits in the Hopper generation. In our table its compute rows match the H100, it carries 96 GB of HBM3 and its interconnect is NVLink-C2C. NVIDIA's H100 page (read 21 September 2026) puts NVLink-C2C at 900 GB/s and calls it "7X faster than PCIe Gen5". Our sources do not describe the CPU side of the GH200, so we do not describe it here.

A GB200 superchip is one Grace CPU with two Blackwell GPUs, and NVIDIA presents it as a rack. Its GB200 NVL72 page (read 21 September 2026) describes 36 Grace CPUs and 72 Blackwell GPUs in a liquid-cooled rack whose "72-GPU NVIDIA NVLink domain ... acts as a single, massive GPU". GB300 NVL72 is the same design with 72 Blackwell Ultra GPUs.

For a renter, the point is that a GB part is a share of a rack-scale system, not a card in a server. The same NVIDIA page counts 2,592 Arm Neoverse V2 CPU cores in the rack. Check what unit the price refers to and that your software stack builds for Arm. Our GPU cluster page covers multi-node and reserved pricing.

Data center, workstation, consumer: what each tier lacks

TierExamplesMemoryNVLinkMIG
Data center, SXMA100, H100, H200, B200HBMFull, through the HGX boardUp to seven instances
Data center, PCIeT4, A10, L4, L40S, A100 PCIe, H100 NVLGDDR6 on T4, A10, L4 and L40S; HBM on the A100 and H100 cardsNone on L40S; bridges only on the HBM cardsNo on L40S; yes on A100 and H100 cards
WorkstationRTX A6000, RTX 6000 Ada, RTX PRO 6000 BlackwellGDDR6 or GDDR7Bridges on Ampere and older cards; none listed for Ada or BlackwellUp to four instances on RTX PRO 6000 Blackwell
ConsumerRTX 3090, RTX 4090, RTX 5090GDDR6X or GDDR7On the RTX 3090 but not the RTX 4090 or RTX 5090, in our tableNot on NVIDIA's supported list

The sources for that table, all read on 21 September 2026: NVIDIA's L40S page has the rows "Multi-Instance GPU (MIG) Support: No" and "NVIDIA NVLink Support: No". The A100 datasheet limits the PCIe card to an NVLink bridge with one other GPU, while the SXM version gets full NVLink through the HGX board. An H100 NVL bridges to a single adjacent card, and an H200 NVL to two or four. NVIDIA's workstation NVLink bridge page lists Ampere cards such as the RTX A6000 and no Ada or Blackwell RTX card. NVIDIA's MIG guide (updated 11 September 2026) lists up to seven instances on A100, H100, H200 and B200, up to four on RTX PRO 6000 Blackwell, and no GeForce cards.

What this means in practice: a model that fits on one GPU runs on any tier, so check the cheaper tiers first. An RTX 4090 or an L40S is a sound choice for single-GPU inference. Once a model has to be split across GPUs, the missing NVLink costs you. vLLM's documentation says that on GPUs without NVLink, naming the L40S, you should use pipeline parallelism instead of tensor parallelism. For multi-GPU training or tensor-parallel serving, rent the SXM tier.

Announced, but not in the rental tables

Vera Rubin is the architecture after Blackwell. NVIDIA's site already lists an HGX Rubin NVL8 board and sixth-generation NVLink, and the TensorRT matrix already has a Vera Rubin row with FP4. The figures are preliminary and NVIDIA's pages disagree with each other: on 21 September 2026 the NVLink page gave 3 TB/s per GPU and the HGX page gave 3.6 TB/s. Treat any Rubin number as provisional. On timing, Nebius said on 16 March 2026 that its Meta contract will be served on Vera Rubin with deliveries from early 2027. Our spec table has no Rubin entry while NVIDIA labels its figures preliminary.

How to place a GPU you have never heard of

  1. Find the architecture. Use the letter or the series number in the decoder table. That gives you the precision ceiling: FP16 on Volta and Turing, BF16 from Ampere, FP8 from Ada and Hopper, FP4 from Blackwell. The full support matrix is in the FP8 article linked above.
  2. Find the tier. SXM data center, PCIe data center, workstation or consumer. That gives you the memory type and whether NVLink and MIG exist.
  3. Read the suffix and the memory size. SXM, PCIe and NVL versions of one chip are different products. Our table lists the H100 at 80 to 94 GB because the NVL card carries 94 GB. If a listing does not say which variant it is, ask.
  4. Check the numbers and the price. Look the GPU up in the spec chart, then in the GPU price list. How to read the throughput and bandwidth rows is the subject of the companion article linked above.

A worked example: "RTX PRO 6000 Blackwell". RTX PRO puts it in the workstation tier; Blackwell means FP4 is accelerated. The tier says GDDR memory, MIG in up to four slices and no NVLink on NVIDIA's product page. So it is a large single-GPU inference card with 96 GB, not a building block for an eight-GPU training node.

The decision rule: for new LLM work, rent Ada, Hopper or newer if you want FP8, and treat Ampere as the floor because of BF16. Use anything older only for small inference jobs or legacy code. If the model spans several GPUs, choose the SXM tier whatever the generation. If it fits on one, choose the cheapest tier whose memory holds it. The best GPU for AI guide applies that rule to live prices.

Sources

Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.

Every source was accessed on 21 September 2026. A date in brackets is the publication or last-update date.

Frequently asked questions

What do the letters in NVIDIA data center GPU names mean?

The letter is the architecture: P for Pascal, V for Volta, T for Turing, A for Ampere, L for Ada Lovelace, H for Hopper and B for Blackwell. GB in front, as in GB200 or GB300, means Grace CPUs packaged with Blackwell GPUs.

What is the order of NVIDIA GPU generations?

Among the architectures we track: Pascal, Volta, Turing, Ampere, then Ada Lovelace and Hopper side by side, then Blackwell and Blackwell Ultra. Vera Rubin is announced as the next one, with preliminary figures only.

What is the difference between SXM, PCIe and NVL?

SXM is a socketed module that sits on an HGX baseboard with full NVLink between the GPUs. PCIe is a standard card in a slot. NVL is a PCIe card that links to neighbouring cards of the same type through NVLink bridges.

Which NVIDIA generation do I need for FP8 or FP4?

FP8 is accelerated from Ada Lovelace and Hopper onward, and FP4 from Blackwell onward. Ampere adds BF16 but has no FP8. On the H100 and L40S, NVIDIA's TensorRT matrix marks FP4 as emulation only.

Can consumer RTX cards do what data center GPUs do?

They share the architecture and its precision support, but not the rest. GeForce cards are not on NVIDIA's MIG supported list, and in our spec table the RTX 4090 and RTX 5090 list PCIe as their only interconnect.

What is a GB200 or GB300?

It is a superchip: Grace CPUs and Blackwell GPUs on one module, normally deployed as an NVL72 rack in which 72 GPUs share one NVLink domain. GB300 uses Blackwell Ultra GPUs in the same rack design.

Related Posts