HBM memory is stacked beside the GPU and connected through very wide buses, a design used on data-center GPUs that SK hynix describes in its August 2019 HBM2E announcement. GDDR is conventional graphics memory used on workstation and consumer cards, with Micron describing its board-mounted layout in its graphics DRAM FAQ, retrieved September 28, 2026. Capacity sets what fits; memory bandwidth sets the token-generation ceiling when decode is memory-bound, as NVIDIA explains in its November 2023 inference guide.
Start with capacity, then ask whether you need more bandwidth. A model that loads successfully has passed only the first test. You still need room for requests and enough speed to meet your response-time target. The useful HBM vs GDDR comparison is between complete GPUs running your intended workload.
Read these seven GPUs before choosing a memory generation
| Spec | H100 | H200 | B200 | MI300X | RTX PRO 6000 | RTX 5090 | L40S |
|---|---|---|---|---|---|---|---|
| Memory type | HBM3 | HBM3e | HBM3e | HBM3 | GDDR7 | GDDR7 | GDDR6 |
| VRAM | 80 to 94 GB | 141 GB | 180 to 192 GB | 192 GB | 96 GB | 32 GB | 48 GB |
| Memory bandwidth | 3,350 GB/s | 4,800 GB/s | 8,000 GB/s | 5,300 GB/s | 1,792 GB/s | 1,792 GB/s | 864 GB/s |
| TDP | 700 W | 700 W | 1,000 W | 750 W | 600 W | 575 W | 350 W |
| Launch year | 2022 | 2024 | 2024 | 2023 | 2025 | 2025 | 2023 |
The site's GPU specification reference groups H100 and MI300X under HBM3, and H200 and B200 under HBM3e. RTX PRO 6000 Blackwell and RTX 5090 use GDDR7; L40S uses GDDR6. These are family-level specifications. Read the exact variant before booking, especially when the capacity column gives a range.
The useful surprise is capacity. RTX PRO 6000 has 96 GB of GDDR7, while the H100 family spans 80 to 94 GB. HBM is not a guarantee of more capacity than every GDDR card. The RTX PRO 6000 guide is worth reading if that capacity lets your job stay on one GPU.
Bandwidth tells a different story. Using the family values, H100 versus L40S is 3.9 times, derived from 3350 ÷ 864. H200 versus RTX PRO 6000 is 2.7 times, derived from 4800 ÷ 1792. B200 versus RTX 5090 is 4.5 times, derived from 8000 ÷ 1792. None is a measured application speedup. For RTX PRO 6000, check the edition's own bandwidth before applying the family ratio. Use the H100 versus L40S comparison to narrow a rental choice, then measure the job that matters to you.
HBM memory and GDDR7 use different routes to bandwidth
SK hynix's August 12, 2019 description explains that HBM stacks DRAM dies vertically and connects them through silicon. It places that memory close to the processor. Micron's graphics DRAM FAQ, retrieved September 28, 2026, describes GDDR components soldered to the processor's circuit board. That is the physical answer to "what is HBM": the stack and its wide connection matter together.
Per-pin speed alone cannot rank these designs. Compare the rate with the interface width, and distinguish one memory device from the whole GPU. The following table attributes each rate and date separately because a product announcement is not a standards announcement.
| Generation | Standard date or explicitly identified product milestone | Per-pin rate and its scope | Interface width |
|---|---|---|---|
| HBM2E | JESD235C document: January 2020; Rambus HBM2E interface: March 3, 2020 | JESD235C adds 2.8 and 3.2 Gb/s bins; Rambus describes 3.2 Gb/s | Rambus: 1024 bits per stack |
| HBM3 | JEDEC JESD238 announcement: January 27, 2022 | JEDEC: up to 6.4 Gb/s | Derived: 16 channels × 64 bits = 1024 bits per stack, using JESD238A scope |
| HBM3E | Micron product specification, retrieved September 28, 2026 | Micron VENDOR CLAIM: greater than 9.2 Gb/s | Micron: 1024 I/Os per stack |
| HBM4 | JEDEC JESD270-4 announcement: April 16, 2025 | JEDEC: up to 8 Gb/s | JEDEC: 2048 bits per stack |
| GDDR6 | Samsung dates its 24 Gb/s product development to 2022 (announced July 19, 2023) | Samsung: 24 Gb/s product grade | Micron: 32 data I/Os per component |
| GDDR6X | Micron product introduction: September 1, 2020; not a JEDEC date | Micron launch products: 19 to 21 Gb/s | Not listed |
| GDDR7 | JEDEC JESD239 announcement: March 5, 2024 | Derived standard ceiling: 48.0 Gb/s; Samsung's July 19, 2023 product: 32 Gb/s | Micron graphics-DRAM interface description: 32 bits per device |
The GDDR7 ceiling is derived from JEDEC's 192 GB/s per device × 8 ÷ 32. It is not a shipping GPU's operating rate. For HBM4 in detail, including each memory maker's announced speeds and the GPUs that use it, see HBM4 explained.
Why decode spends its time moving bytes
NVIDIA's February 1, 2023 performance guide defines arithmetic intensity as operations divided by bytes accessed. Its compute and transfer limits give this derived roofline expression:
Achievable operations/s ≤ min(peak operations/s, memory bandwidth × operations/byte).
When an operation does little work per byte, faster arithmetic cannot remove the transfer limit. NVIDIA's November 17, 2023 inference guide explains that moving weights, keys, values and activations can dominate decode latency. The same guide distinguishes prefill, which processes the known prompt, from decode, which generates subsequent tokens one at a time.
For a deliberately simplified single-stream example, assume a dense model with 70 billion parameters stored at one byte each. NVIDIA's September 1, 2026 sizing guidance gives that byte count for FP8 weights. The raw weights are 70 GB, derived. Assume every token reads those weights once from memory, with no weight reuse between decode steps.
Theoretical upper bound, not a measurement: tokens/s per stream ≤ bandwidth ÷ bytes read per token.
Using H200's family bandwidth and only those weight bytes gives 4800 GB/s ÷ 70 GB = 68.6 tokens/s, derived. This optimistic ceiling ignores cache traffic and other work. It is not a forecast for a named model or serving engine, and it does not describe aggregate batched throughput.
NVIDIA's November 2023 guide explains that batching shares weight reads across requests. The Scaling Book's inference chapter, retrieved September 28, 2026, adds the crucial limit: sufficiently large batches can make decode's feed-forward operations compute-bound, while each request still has its own attention cache. More bandwidth helps the memory-limited parts; it does not make every stage faster by the bandwidth ratio.
Use how to read GPU specs for AI to separate these limits. For workload differences, see training versus inference.
Capacity means weights plus the running requests
The Scaling Book gives the cache allocation for one sequence as:
KV cache bytes = 2 × bytes per value × H × K × L × T.
Here H is head dimension, K is the number of KV heads, L is the layer count and T is the cached token count. The factor of two represents keys and values. Multiplying by B gives the derived allocation for B independent, equal-length sequences. Use KV heads, rather than substituting query heads.
NVIDIA's November 2023 guide identifies linear cache growth with context length and batch size. This is why a successful short prompt is a poor capacity test for a service handling long conversations. Reserve memory for the intended request load before deciding that the model fits.
Start with the LLM VRAM calculator, then inspect the serving engine's actual allocation. Set the model format, context limit and concurrency first. Compare GPUs after those choices, rather than shrinking the workload silently to suit the cheapest card.
What the HBM premium buys
The live table below shows the rental comparison for the main candidates.
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| H100 | $2.50 | Hyperstack | 9 |
| H200 | $3.43 | QuantaCloud | 6 |
| B200 | $6.79 | RunPod | 3 |
| RTX PRO 6000 Blackwell | $0.59 | RunPod | 4 |
| RTX 5090 | $0.53 | Vast.ai | 3 |
| L40S | $0.97 | Massed Compute | 5 |
Where an HBM option costs more, the price gap buys the capacity and bandwidth differences shown above. It earns its place only if those differences let your intended job fit or meet its performance target. Do not pay for a memory label by itself.
HBM also carries component and packaging costs. ANALYST ESTIMATE: Semiconductor Engineering reported TechInsights' roughly US$120 estimate for a 16 GB HBM2 component on December 17, 2019, excluding packaging. That is historical component pricing, not a current HBM3E quote. ANALYST ESTIMATE: TrendForce's October 30, 2025 assessment put HBM3e price increases during 2025 at 5 to 10%.
Packaging can constrain supply separately from the memory dies. TSMC CEO C.C. Wei said on January 16, 2025 that CoWoS capacity was very tight and could not meet customer needs. That dated statement does not establish the severity of a September 2026 bottleneck.
ANALYST ESTIMATE: KB Securities' September 25, 2025 report, citing Counterpoint, assigned Q2 2025 HBM market shares of 62% to SK hynix, 21% to Micron and 17% to Samsung. These are HBM market shares, with no revenue, bit or unit denominator specified. Do not interpret them as shares of all DRAM or as current supplier inventory.
Unified memory is a separate capacity choice
Apple's October 30, 2024 announcement specified 120 GB/s for M4, 273 GB/s for M4 Pro and up to 546 GB/s for M4 Max. Its March 5, 2025 M3 Ultra announcement specified over 800 GB/s. Those are dated generation examples, not a claim about Apple's current lineup.
NVIDIA's DGX Spark specification, retrieved September 28, 2026, lists 128 GB of coherent unified LPDDR5x memory and 273 GB/s bandwidth for its GB10 system. AMD's Ryzen AI Max+ 395 system specification, retrieved the same day, lists 128 GB of LPDDR5x and 256 GB/s. AMD's Max PRO whitepaper, retrieved the same day, also distinguishes the system pool from graphics allocation: up to 96 GB of 128 GB can be dedicated to graphics.
Compare usable capacity and bandwidth independently. A large shared pool answers a different question from how quickly the GPU can read it. The DGX Spark versus cloud GPU guide develops that desktop decision.
Every GPU family, grouped by memory type
This inventory uses the site's family specifications throughout. Capacity ranges show the recorded minimum and maximum, while bandwidth is the recorded family value. Do not pair a range endpoint with that bandwidth and assume it describes every variant. Older memory types remain here so you can place an unfamiliar rental listing.
| Memory | GPU family | Capacity (GB) | Bandwidth (GB/s) |
|---|---|---|---|
| HBM2 | A30 | 24 | 933 |
| HBM2 | P100 | 16 | 732 |
| HBM2 | TITAN V | 12 | 653 |
| HBM2 | V100 | 16 to 32 | 900 |
| HBM2e | A100 | 40 to 80 | 2039 |
| HBM2e | Gaudi 2 | 96 | 2450 |
| HBM2e | MI250X | 128 | 3277 |
| HBM3 | GH200 | 96 | 4000 |
| HBM3 | H100 | 80 to 94 | 3350 |
| HBM3 | MI300X | 192 | 5300 |
| HBM3e | B200 | 180 to 192 | 8000 |
| HBM3e | B300 | 262 to 288 | 8000 |
| HBM3e | GB300 | 262 to 288 | 8000 |
| HBM3e | H200 | 141 | 4800 |
| HBM3e | MI325X | 256 | 6000 |
| HBM3e | MI355X | 288 | 8000 |
| GDDR5 | GTX 1070 | 8 | 256 |
| GDDR5 | Quadro M4000 | 8 | 192 |
| GDDR5 | Quadro P4000 | 8 | 243 |
| GDDR5X | GTX 1080 | 8 to 11 | 320 |
| GDDR5X | Quadro P5000 | 16 | 288 |
| GDDR5X | Quadro P6000 | 24 | 432 |
| GDDR5X | TITAN Xp | 12 | 548 |
| GDDR6 | A10 | 24 | 600 |
| GDDR6 | A16 | 16 | 200 |
| GDDR6 | A40 | 48 | 696 |
| GDDR6 | L4 | 24 | 300 |
| GDDR6 | L40 | 48 | 864 |
| GDDR6 | L40S | 48 | 864 |
| GDDR6 | Quadro RTX 4000 | 8 | 416 |
| GDDR6 | Quadro RTX 5000 | 16 | 448 |
| GDDR6 | Quadro RTX 6000 | 24 | 672 |
| GDDR6 | Quadro RTX 8000 | 48 | 672 |
| GDDR6 | RTX 2000 Ada | 16 | 224 |
| GDDR6 | RTX 2060 | 6 to 12 | 336 |
| GDDR6 | RTX 2070 | 8 | 448 |
| GDDR6 | RTX 2080 | 8 to 11 | 448 |
| GDDR6 | RTX 3060 | 8 to 12 | 360 |
| GDDR6 | RTX 3070 | 8 | 448 |
| GDDR6 | RTX 4000 Ada | 20 | 360 |
| GDDR6 | RTX 4060 | 8 to 16 | 272 |
| GDDR6 | RTX 4500 Ada | 24 | 432 |
| GDDR6 | RTX 5000 Ada | 32 | 576 |
| GDDR6 | RTX 5880 Ada | 48 | 960 |
| GDDR6 | RTX 6000 Ada | 48 | 960 |
| GDDR6 | RTX A2000 | 6 to 12 | 288 |
| GDDR6 | RTX A4000 | 16 to 20 | 448 |
| GDDR6 | RTX A5000 | 24 | 768 |
| GDDR6 | RTX A6000 | 48 | 768 |
| GDDR6 | T4 | 16 | 320 |
| GDDR6X | RTX 3080 | 10 to 12 | 760 |
| GDDR6X | RTX 3090 | 24 | 936 |
| GDDR6X | RTX 4070 | 12 to 16 | 504 |
| GDDR6X | RTX 4080 | 16 | 717 |
| GDDR6X | RTX 4090 | 24 | 1008 |
| GDDR7 | RTX 5060 | 8 to 16 | 448 |
| GDDR7 | RTX 5070 | 12 to 16 | 672 |
| GDDR7 | RTX 5080 | 16 | 960 |
| GDDR7 | RTX 5090 | 32 | 1792 |
| GDDR7 | RTX PRO 6000 | 96 | 1792 |
The live table below covers the remaining families, including MI300X. Match the family name to the inventory before comparing prices.
Choose GDDR unless capacity or measured latency rules it out
Rent a GDDR card when your full allocation fits and it meets your decode latency and concurrency targets. Pay for HBM when the GDDR candidates fail one of those requirements and the HBM candidate passes. Keep the model, precision, context and request load fixed during that comparison. Choose the least expensive configuration that passes, then increase capacity or bandwidth only for a demonstrated constraint.
Sources
- SK hynix: HBM2E architecture, 2019-08-12.
- Micron: graphics DRAM FAQ, retrieved 2026-09-28.
- Micron: graphics memory interface, retrieved 2026-09-28.
- JESD235C document, retrieved 2026-09-28.
- Rambus: HBM2E interface, 2020-03-03.
- JEDEC: HBM3 announcement, 2022-01-27.
- JESD238A standard scope, 2023-01-01.
- Micron: HBM3E product specification, retrieved 2026-09-28.
- JEDEC: HBM4 announcement, 2025-04-16.
- Samsung: GDDR6 and GDDR7 milestones, 2023-07-19.
- Micron: GDDR6X introduction, 2020-09-01.
- JEDEC: GDDR7 announcement, 2024-03-05.
- NVIDIA: GPU performance background, 2023-02-01.
- NVIDIA: inference optimization, 2023-11-17.
- The Scaling Book: inference, retrieved 2026-09-28.
- NVIDIA: inference sizing, 2026-09-01.
- Semiconductor Engineering: TechInsights HBM estimate, 2019-12-17.
- TrendForce: HBM3e price assessment, 2025-10-30.
- TSMC: fourth-quarter 2024 earnings transcript, 2025-01-16.
- KB Securities citing Counterpoint: HBM market shares, 2025-09-25.
- Apple: M4 family bandwidth, 2024-10-30.
- Apple: M3 Ultra announcement, 2025-03-05.
- NVIDIA: DGX Spark specifications, retrieved 2026-09-28.
- AMD: Ryzen AI Max+ 395 system, retrieved 2026-09-28.
- AMD: Ryzen AI Max PRO whitepaper, retrieved 2026-09-28.