An NVIDIA data center GPU name is a letter for the architecture, a number for the tier inside it, and a suffix for the form factor. A is Ampere, L is Ada Lovelace, H is Hopper and B is Blackwell. GB in front means Grace CPUs are packaged with Blackwell GPUs. Once you know the generation and the tier, you know which number formats the card accelerates, whether it has NVLink and MIG, and roughly where it will sit in a price table.
The decoder table
| Letter or suffix | Meaning | Example |
|---|---|---|
| P | Pascal data center part | P100 |
| V | Volta | V100 |
| T | Turing | T4 |
| A | Ampere | A100, A40, A10 |
| L | Ada Lovelace data center part | L4, L40S |
| H | Hopper | H100, H200 |
| B | Blackwell, and Blackwell Ultra for B300 | B200, B300 |
| GB | Grace CPUs packaged with Blackwell GPUs (a superchip) | GB200, GB300 |
| GH200 | Hopper-generation part; our table lists its interconnect as NVLink-C2C | GH200 |
| GeForce RTX 20, 30, 40, 50 series | Consumer cards: Turing, Ampere, Ada Lovelace, Blackwell in that order | RTX 3090, RTX 4090, RTX 5090 |
| GTX | Older consumer cards; the Pascal ones we track have no Tensor cores | GTX 1080 |
| RTX A plus a number | Ampere workstation card | RTX A6000 |
| RTX, a number, then "Ada" | Ada Lovelace workstation card | RTX 6000 Ada |
| RTX PRO | Blackwell workstation and server card | RTX PRO 6000 Blackwell |
| Quadro | Workstation brand; in our table it runs from Maxwell to Turing | Quadro RTX 8000 |
| SXM | Socketed module on an HGX baseboard, full NVLink | H100 SXM5 |
| PCIe | Standard add-in card | A100 PCIe |
| NVL (on a card) | PCIe card joined to neighbours by NVLink bridges | H100 NVL, H200 NVL |
| NVL8, NVL72 | Number of GPUs in one NVLink domain | GB200 NVL72 |
| HGX, DGX | HGX is the eight-GPU baseboard sold to server makers; DGX is NVIDIA's own complete system built on it | HGX B200, DGX B200 |
Two notes on SXM. NVIDIA does not publish what the letters stand for, so we do not expand it. The module is numbered by generation: Wikipedia's SXM article (read 21 September 2026) pairs SXM2 with P100 and V100, SXM4 with A100, SXM5 with H100 and H200, and SXM6 with B200. NVIDIA's own pages say only "Blackwell SXM"; the SXM6 label comes from server makers such as Dell and Lenovo.
The number after the letter ranks cards inside a generation only loosely. An L4 is a 72 W card with 24 GB and an L40S is a 350 W card with 48 GB, but an H200 is not twice an H100. It is the same GPU with more and faster memory, which we cover in how to read GPU specs for AI.
NVIDIA data center GPU list by generation
This list is built from the launch years and architectures in our datasheet-verified spec table. It covers the GPUs this site tracks, not everything NVIDIA has shipped. The full GPU spec chart has every row.
| Architecture | Launch years in our table | GPUs we track |
|---|---|---|
| Pascal | 2016 to 2017 | P100, GTX 1070, GTX 1080, Quadro P4000, P5000 and P6000, TITAN Xp |
| Volta | 2017 | V100, TITAN V |
| Turing | 2018 to 2019 | T4, Quadro RTX 4000, 5000, 6000 and 8000, RTX 2060, 2070 and 2080 |
| Ampere | 2020 to 2021 | A100, A40, A30, A10, A16, RTX A2000 to A6000, RTX 3060 to 3090 |
| Ada Lovelace | 2022 to 2024 | L4, L40, L40S, RTX 2000 Ada to RTX 6000 Ada, RTX 4060 to 4090 |
| Hopper | 2022 to 2024 | H100 (2022), GH200 (2023), H200 (2024) |
| Blackwell | 2024 to 2025 | B200 (2024), RTX 5060 to 5090, RTX PRO 6000 Blackwell (2025) |
| Blackwell Ultra | 2025 | B300, GB300 |
What each generation rents for now
What changed each generation, for a renter
| Spec | V100 | A100 | H100 | B200 |
|---|---|---|---|---|
| Launch year | 2017 | 2020 | 2022 | 2024 |
| Architecture | Volta | Ampere | Hopper | Blackwell |
| VRAM | 16 to 32 GB | 40 to 80 GB | 80 to 94 GB | 180 to 192 GB |
| Memory bandwidth | 900 GB/s | 2,039 GB/s | 3,350 GB/s | 8,000 GB/s |
| FP16 (dense) | 125 TFLOPS | 312 TFLOPS | 989 TFLOPS | 2,250 TFLOPS |
| FP16 (with sparsity) | Not published | 624 TFLOPS | 1,979 TFLOPS | 4,500 TFLOPS |
| FP8 (dense) | Not published | Not published | 1,979 TFLOPS | 4,500 TFLOPS |
| FP8 (with sparsity) | Not published | Not published | 3,958 TFLOPS | 9,000 TFLOPS |
| FP4 (dense) | Not published | Not published | Not published | 9,000 TFLOPS |
| FP4 (with sparsity) | Not published | Not published | Not published | 18,000 TFLOPS |
Pascal (2016). No Tensor cores. The P100 already has the shape of the later data center parts: HBM2 memory, NVLink and an SXM module. Treat Pascal as legacy.
Volta (2017). The first Tensor cores: the V100 has 640 of them beside 5,120 CUDA cores, and with them came FP16 mixed-precision training. It has 16 or 32 GB of HBM2 and second-generation NVLink, which NVIDIA's V100 datasheet lists at 300 GB/s on the SXM2 module. Software support is thinning. NVIDIA's TensorRT support matrix (8 September 2026) no longer lists Volta at all.
Turing (2018). Adds INT8 and INT4 to the Tensor cores, according to SemiAnalysis's history of the design (23 June 2025). The T4, a 70 W card with 16 GB of GDDR6, is the Turing data center part we track. The TensorRT matrix gives it FP32, FP16 and INT8 only, and marks BF16 as not available.
Ampere (2020). Three things arrive at once: BF16 and TF32, MIG partitioning, and third-generation NVLink at 600 GB/s per GPU on the A100 datasheet. BF16 is the one that matters most for training, for reasons covered in FP8 vs FP16 vs BF16. Ampere has no FP8. It is the oldest generation we would choose for new training work. The A100 vs H100 comparison shows what the step up costs today.
Ada Lovelace (2022) and Hopper (2022). These two are the same age and both add FP8; NVIDIA's Transformer Engine needs compute capability 8.9 or higher, which means Ada, Hopper or Blackwell. They split by memory and interconnect. Hopper is the HBM part, with 80 GB of HBM3 on the H100, fourth-generation NVLink at 900 GB/s per GPU and PCIe Gen5, followed by 141 GB of HBM3e on the H200 two years later. Ada is the PCIe part: GDDR6 memory on a PCIe Gen4 card, and on the L40S a block of RT cores for graphics beside the Tensor cores. On the H100 and the L40S, the TensorRT matrix marks FP4 as "hardware emulation mode", so the hardware does not accelerate it.
Blackwell (2024) and Blackwell Ultra (2025). Adds FP4. SemiAnalysis lists NVIDIA's own NVFP4 format and the MXFP8, MXFP6 and MXFP4 microscaling formats as new in this generation. It also brings fifth-generation NVLink at 1,800 GB/s per GPU on NVIDIA's NVLink page (read 21 September 2026). The B200 carries 180 to 192 GB of HBM3e at 8,000 GB/s and a 1,000 W TDP. FP4 also reaches the consumer and workstation cards: NVIDIA lists it for both the RTX 5090 and the RTX PRO 6000 Blackwell, which use GDDR7. B300 is the Blackwell Ultra part, at 1,400 W. Its memory is quoted two ways: on 21 September 2026 Dell listed 270 GB per GPU and Latitude.sh listed 288 GB. No source we found explains the gap, so check the figure on the listing you rent. The B200 vs H200 comparison tracks the price gap between the two newest generations.
All the NVLink figures above come from NVIDIA's datasheets and NVLink page as read on 21 September 2026, and they are totals for both directions added together. Per direction is half, which is why Lenovo's B200 page (updated 13 January 2026) lists 900 GB/s.
Superchips: GH200, GB200 and GB300
The GH200 sits in the Hopper generation. In our table its compute rows match the H100, it carries 96 GB of HBM3 and its interconnect is NVLink-C2C. NVIDIA's H100 page (read 21 September 2026) puts NVLink-C2C at 900 GB/s and calls it "7X faster than PCIe Gen5". Our sources do not describe the CPU side of the GH200, so we do not describe it here.
A GB200 superchip is one Grace CPU with two Blackwell GPUs, and NVIDIA presents it as a rack. Its GB200 NVL72 page (read 21 September 2026) describes 36 Grace CPUs and 72 Blackwell GPUs in a liquid-cooled rack whose "72-GPU NVIDIA NVLink domain ... acts as a single, massive GPU". GB300 NVL72 is the same design with 72 Blackwell Ultra GPUs.
For a renter, the point is that a GB part is a share of a rack-scale system, not a card in a server. The same NVIDIA page counts 2,592 Arm Neoverse V2 CPU cores in the rack. Check what unit the price refers to and that your software stack builds for Arm. Our GPU cluster page covers multi-node and reserved pricing.
Data center, workstation, consumer: what each tier lacks
| Tier | Examples | Memory | NVLink | MIG |
|---|---|---|---|---|
| Data center, SXM | A100, H100, H200, B200 | HBM | Full, through the HGX board | Up to seven instances |
| Data center, PCIe | T4, A10, L4, L40S, A100 PCIe, H100 NVL | GDDR6 on T4, A10, L4 and L40S; HBM on the A100 and H100 cards | None on L40S; bridges only on the HBM cards | No on L40S; yes on A100 and H100 cards |
| Workstation | RTX A6000, RTX 6000 Ada, RTX PRO 6000 Blackwell | GDDR6 or GDDR7 | Bridges on Ampere and older cards; none listed for Ada or Blackwell | Up to four instances on RTX PRO 6000 Blackwell |
| Consumer | RTX 3090, RTX 4090, RTX 5090 | GDDR6X or GDDR7 | On the RTX 3090 but not the RTX 4090 or RTX 5090, in our table | Not on NVIDIA's supported list |
The sources for that table, all read on 21 September 2026: NVIDIA's L40S page has the rows "Multi-Instance GPU (MIG) Support: No" and "NVIDIA NVLink Support: No". The A100 datasheet limits the PCIe card to an NVLink bridge with one other GPU, while the SXM version gets full NVLink through the HGX board. An H100 NVL bridges to a single adjacent card, and an H200 NVL to two or four. NVIDIA's workstation NVLink bridge page lists Ampere cards such as the RTX A6000 and no Ada or Blackwell RTX card. NVIDIA's MIG guide (updated 11 September 2026) lists up to seven instances on A100, H100, H200 and B200, up to four on RTX PRO 6000 Blackwell, and no GeForce cards.
What this means in practice: a model that fits on one GPU runs on any tier, so check the cheaper tiers first. An RTX 4090 or an L40S is a sound choice for single-GPU inference. Once a model has to be split across GPUs, the missing NVLink costs you. vLLM's documentation says that on GPUs without NVLink, naming the L40S, you should use pipeline parallelism instead of tensor parallelism. For multi-GPU training or tensor-parallel serving, rent the SXM tier.
Announced, but not in the rental tables
Vera Rubin is the architecture after Blackwell. NVIDIA's site already lists an HGX Rubin NVL8 board and sixth-generation NVLink, and the TensorRT matrix already has a Vera Rubin row with FP4. The figures are preliminary and NVIDIA's pages disagree with each other: on 21 September 2026 the NVLink page gave 3 TB/s per GPU and the HGX page gave 3.6 TB/s. Treat any Rubin number as provisional. On timing, Nebius said on 16 March 2026 that its Meta contract will be served on Vera Rubin with deliveries from early 2027. Our spec table has no Rubin entry while NVIDIA labels its figures preliminary.
How to place a GPU you have never heard of
- Find the architecture. Use the letter or the series number in the decoder table. That gives you the precision ceiling: FP16 on Volta and Turing, BF16 from Ampere, FP8 from Ada and Hopper, FP4 from Blackwell. The full support matrix is in the FP8 article linked above.
- Find the tier. SXM data center, PCIe data center, workstation or consumer. That gives you the memory type and whether NVLink and MIG exist.
- Read the suffix and the memory size. SXM, PCIe and NVL versions of one chip are different products. Our table lists the H100 at 80 to 94 GB because the NVL card carries 94 GB. If a listing does not say which variant it is, ask.
- Check the numbers and the price. Look the GPU up in the spec chart, then in the GPU price list. How to read the throughput and bandwidth rows is the subject of the companion article linked above.
A worked example: "RTX PRO 6000 Blackwell". RTX PRO puts it in the workstation tier; Blackwell means FP4 is accelerated. The tier says GDDR memory, MIG in up to four slices and no NVLink on NVIDIA's product page. So it is a large single-GPU inference card with 96 GB, not a building block for an eight-GPU training node.
The decision rule: for new LLM work, rent Ada, Hopper or newer if you want FP8, and treat Ampere as the floor because of BF16. Use anything older only for small inference jobs or legacy code. If the model spans several GPUs, choose the SXM tier whatever the generation. If it fits on one, choose the cheapest tier whose memory holds it. The best GPU for AI guide applies that rule to live prices.
Sources
Source pages were retrieved with automated research tools on the dates shown and each figure was traced back to its source before publishing. Rental prices on this page are not typed: they are read live from GPUperhour's own data.
Every source was accessed on 21 September 2026. A date in brackets is the publication or last-update date.
- NVIDIA NVLink page
- NVIDIA HGX platform page
- NVIDIA DGX B200 page
- NVIDIA GB200 NVL72 page
- NVIDIA H100 product page and H100 NVL product brief (PDF) (March 2024)
- NVIDIA H200 product page
- NVIDIA L40S product page
- NVIDIA A100 datasheet (PDF)
- NVIDIA V100 datasheet (PDF)
- NVIDIA GeForce RTX 5090 product page
- NVIDIA RTX PRO 6000 Blackwell product page
- NVIDIA workstation NVLink bridges page
- NVIDIA MIG user guide: supported GPUs (updated 11 September 2026)
- NVIDIA TensorRT support matrix (8 September 2026)
- NVIDIA Transformer Engine repository
- SemiAnalysis: NVIDIA Tensor Core Evolution, from Volta to Blackwell (23 June 2025)
- Wikipedia: SXM (socket)
- Dell PowerEdge XE9780 configuration page
- Lenovo Press: ThinkSystem NVIDIA B200 180GB 1000W GPU (updated 13 January 2026)
- Latitude.sh pricing page
- vLLM docs: parallelism and scaling
- Nebius: AI infrastructure agreement with Meta (16 March 2026)