HBM vs GDDR: Pay for Bandwidth After the Model Fits

Compare HBM memory and GDDR7 by capacity, bandwidth and live GPU rental prices. Learn when faster memory helps LLM decode and when a GDDR card is enough.

By Faiz Ahmed•
•13 min read

HBM memory is stacked beside the GPU and connected through very wide buses, a design used on data-center GPUs that SK hynix describes in its August 2019 HBM2E announcement. GDDR is conventional graphics memory used on workstation and consumer cards, with Micron describing its board-mounted layout in its graphics DRAM FAQ, retrieved September 28, 2026. Capacity sets what fits; memory bandwidth sets the token-generation ceiling when decode is memory-bound, as NVIDIA explains in its November 2023 inference guide.

Start with capacity, then ask whether you need more bandwidth. A model that loads successfully has passed only the first test. You still need room for requests and enough speed to meet your response-time target. The useful HBM vs GDDR comparison is between complete GPUs running your intended workload.

Read these seven GPUs before choosing a memory generation

SpecH100H200B200MI300XRTX PRO 6000RTX 5090L40S
Memory typeHBM3HBM3eHBM3eHBM3GDDR7GDDR7GDDR6
VRAM80 to 94 GB141 GB180 to 192 GB192 GB96 GB32 GB48 GB
Memory bandwidth3,350 GB/s4,800 GB/s8,000 GB/s5,300 GB/s1,792 GB/s1,792 GB/s864 GB/s
TDP700 W700 W1,000 W750 W600 W575 W350 W
Launch year2022202420242023202520252023
Figures from the vendor datasheets: H100, H200, B200, MI300X, RTX PRO 6000, RTX 5090, L40S, checked 13 Sep 2026. "Not published" means the vendor gives no figure.

The site's GPU specification reference groups H100 and MI300X under HBM3, and H200 and B200 under HBM3e. RTX PRO 6000 Blackwell and RTX 5090 use GDDR7; L40S uses GDDR6. These are family-level specifications. Read the exact variant before booking, especially when the capacity column gives a range.

The useful surprise is capacity. RTX PRO 6000 has 96 GB of GDDR7, while the H100 family spans 80 to 94 GB. HBM is not a guarantee of more capacity than every GDDR card. The RTX PRO 6000 guide is worth reading if that capacity lets your job stay on one GPU.

Bandwidth tells a different story. Using the family values, H100 versus L40S is 3.9 times, derived from 3350 ÷ 864. H200 versus RTX PRO 6000 is 2.7 times, derived from 4800 ÷ 1792. B200 versus RTX 5090 is 4.5 times, derived from 8000 ÷ 1792. None is a measured application speedup. For RTX PRO 6000, check the edition's own bandwidth before applying the family ratio. Use the H100 versus L40S comparison to narrow a rental choice, then measure the job that matters to you.

HBM memory and GDDR7 use different routes to bandwidth

SK hynix's August 12, 2019 description explains that HBM stacks DRAM dies vertically and connects them through silicon. It places that memory close to the processor. Micron's graphics DRAM FAQ, retrieved September 28, 2026, describes GDDR components soldered to the processor's circuit board. That is the physical answer to "what is HBM": the stack and its wide connection matter together.

Per-pin speed alone cannot rank these designs. Compare the rate with the interface width, and distinguish one memory device from the whole GPU. The following table attributes each rate and date separately because a product announcement is not a standards announcement.

GenerationStandard date or explicitly identified product milestonePer-pin rate and its scopeInterface width
HBM2EJESD235C document: January 2020; Rambus HBM2E interface: March 3, 2020JESD235C adds 2.8 and 3.2 Gb/s bins; Rambus describes 3.2 Gb/sRambus: 1024 bits per stack
HBM3JEDEC JESD238 announcement: January 27, 2022JEDEC: up to 6.4 Gb/sDerived: 16 channels × 64 bits = 1024 bits per stack, using JESD238A scope
HBM3EMicron product specification, retrieved September 28, 2026Micron VENDOR CLAIM: greater than 9.2 Gb/sMicron: 1024 I/Os per stack
HBM4JEDEC JESD270-4 announcement: April 16, 2025JEDEC: up to 8 Gb/sJEDEC: 2048 bits per stack
GDDR6Samsung dates its 24 Gb/s product development to 2022 (announced July 19, 2023)Samsung: 24 Gb/s product gradeMicron: 32 data I/Os per component
GDDR6XMicron product introduction: September 1, 2020; not a JEDEC dateMicron launch products: 19 to 21 Gb/sNot listed
GDDR7JEDEC JESD239 announcement: March 5, 2024Derived standard ceiling: 48.0 Gb/s; Samsung's July 19, 2023 product: 32 Gb/sMicron graphics-DRAM interface description: 32 bits per device

The GDDR7 ceiling is derived from JEDEC's 192 GB/s per device × 8 ÷ 32. It is not a shipping GPU's operating rate. For HBM4 in detail, including each memory maker's announced speeds and the GPUs that use it, see HBM4 explained.

Why decode spends its time moving bytes

NVIDIA's February 1, 2023 performance guide defines arithmetic intensity as operations divided by bytes accessed. Its compute and transfer limits give this derived roofline expression:

Achievable operations/s ≤ min(peak operations/s, memory bandwidth × operations/byte).

When an operation does little work per byte, faster arithmetic cannot remove the transfer limit. NVIDIA's November 17, 2023 inference guide explains that moving weights, keys, values and activations can dominate decode latency. The same guide distinguishes prefill, which processes the known prompt, from decode, which generates subsequent tokens one at a time.

For a deliberately simplified single-stream example, assume a dense model with 70 billion parameters stored at one byte each. NVIDIA's September 1, 2026 sizing guidance gives that byte count for FP8 weights. The raw weights are 70 GB, derived. Assume every token reads those weights once from memory, with no weight reuse between decode steps.

Theoretical upper bound, not a measurement: tokens/s per stream ≤ bandwidth ÷ bytes read per token.

Using H200's family bandwidth and only those weight bytes gives 4800 GB/s ÷ 70 GB = 68.6 tokens/s, derived. This optimistic ceiling ignores cache traffic and other work. It is not a forecast for a named model or serving engine, and it does not describe aggregate batched throughput.

NVIDIA's November 2023 guide explains that batching shares weight reads across requests. The Scaling Book's inference chapter, retrieved September 28, 2026, adds the crucial limit: sufficiently large batches can make decode's feed-forward operations compute-bound, while each request still has its own attention cache. More bandwidth helps the memory-limited parts; it does not make every stage faster by the bandwidth ratio.

Use how to read GPU specs for AI to separate these limits. For workload differences, see training versus inference.

Capacity means weights plus the running requests

The Scaling Book gives the cache allocation for one sequence as:

KV cache bytes = 2 × bytes per value × H × K × L × T.

Here H is head dimension, K is the number of KV heads, L is the layer count and T is the cached token count. The factor of two represents keys and values. Multiplying by B gives the derived allocation for B independent, equal-length sequences. Use KV heads, rather than substituting query heads.

NVIDIA's November 2023 guide identifies linear cache growth with context length and batch size. This is why a successful short prompt is a poor capacity test for a service handling long conversations. Reserve memory for the intended request load before deciding that the model fits.

Start with the LLM VRAM calculator, then inspect the serving engine's actual allocation. Set the model format, context limit and concurrency first. Compare GPUs after those choices, rather than shrinking the workload silently to suit the cheapest card.

What the HBM premium buys

The live table below shows the rental comparison for the main candidates.

GPUCheapest $/GPU-hrProviderProviders in stock
H100$2.50Hyperstack9
H200$3.43QuantaCloud6
B200$6.79RunPod3
RTX PRO 6000 Blackwell$0.59RunPod4
RTX 5090$0.53Vast.ai3
L40S$0.97Massed Compute5
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Where an HBM option costs more, the price gap buys the capacity and bandwidth differences shown above. It earns its place only if those differences let your intended job fit or meet its performance target. Do not pay for a memory label by itself.

HBM also carries component and packaging costs. ANALYST ESTIMATE: Semiconductor Engineering reported TechInsights' roughly US$120 estimate for a 16 GB HBM2 component on December 17, 2019, excluding packaging. That is historical component pricing, not a current HBM3E quote. ANALYST ESTIMATE: TrendForce's October 30, 2025 assessment put HBM3e price increases during 2025 at 5 to 10%.

Packaging can constrain supply separately from the memory dies. TSMC CEO C.C. Wei said on January 16, 2025 that CoWoS capacity was very tight and could not meet customer needs. That dated statement does not establish the severity of a September 2026 bottleneck.

ANALYST ESTIMATE: KB Securities' September 25, 2025 report, citing Counterpoint, assigned Q2 2025 HBM market shares of 62% to SK hynix, 21% to Micron and 17% to Samsung. These are HBM market shares, with no revenue, bit or unit denominator specified. Do not interpret them as shares of all DRAM or as current supplier inventory.

Unified memory is a separate capacity choice

Apple's October 30, 2024 announcement specified 120 GB/s for M4, 273 GB/s for M4 Pro and up to 546 GB/s for M4 Max. Its March 5, 2025 M3 Ultra announcement specified over 800 GB/s. Those are dated generation examples, not a claim about Apple's current lineup.

NVIDIA's DGX Spark specification, retrieved September 28, 2026, lists 128 GB of coherent unified LPDDR5x memory and 273 GB/s bandwidth for its GB10 system. AMD's Ryzen AI Max+ 395 system specification, retrieved the same day, lists 128 GB of LPDDR5x and 256 GB/s. AMD's Max PRO whitepaper, retrieved the same day, also distinguishes the system pool from graphics allocation: up to 96 GB of 128 GB can be dedicated to graphics.

Compare usable capacity and bandwidth independently. A large shared pool answers a different question from how quickly the GPU can read it. The DGX Spark versus cloud GPU guide develops that desktop decision.

Every GPU family, grouped by memory type

This inventory uses the site's family specifications throughout. Capacity ranges show the recorded minimum and maximum, while bandwidth is the recorded family value. Do not pair a range endpoint with that bandwidth and assume it describes every variant. Older memory types remain here so you can place an unfamiliar rental listing.

MemoryGPU familyCapacity (GB)Bandwidth (GB/s)
HBM2A3024933
HBM2P10016732
HBM2TITAN V12653
HBM2V10016 to 32900
HBM2eA10040 to 802039
HBM2eGaudi 2962450
HBM2eMI250X1283277
HBM3GH200964000
HBM3H10080 to 943350
HBM3MI300X1925300
HBM3eB200180 to 1928000
HBM3eB300262 to 2888000
HBM3eGB300262 to 2888000
HBM3eH2001414800
HBM3eMI325X2566000
HBM3eMI355X2888000
GDDR5GTX 10708256
GDDR5Quadro M40008192
GDDR5Quadro P40008243
GDDR5XGTX 10808 to 11320
GDDR5XQuadro P500016288
GDDR5XQuadro P600024432
GDDR5XTITAN Xp12548
GDDR6A1024600
GDDR6A1616200
GDDR6A4048696
GDDR6L424300
GDDR6L4048864
GDDR6L40S48864
GDDR6Quadro RTX 40008416
GDDR6Quadro RTX 500016448
GDDR6Quadro RTX 600024672
GDDR6Quadro RTX 800048672
GDDR6RTX 2000 Ada16224
GDDR6RTX 20606 to 12336
GDDR6RTX 20708448
GDDR6RTX 20808 to 11448
GDDR6RTX 30608 to 12360
GDDR6RTX 30708448
GDDR6RTX 4000 Ada20360
GDDR6RTX 40608 to 16272
GDDR6RTX 4500 Ada24432
GDDR6RTX 5000 Ada32576
GDDR6RTX 5880 Ada48960
GDDR6RTX 6000 Ada48960
GDDR6RTX A20006 to 12288
GDDR6RTX A400016 to 20448
GDDR6RTX A500024768
GDDR6RTX A600048768
GDDR6T416320
GDDR6XRTX 308010 to 12760
GDDR6XRTX 309024936
GDDR6XRTX 407012 to 16504
GDDR6XRTX 408016717
GDDR6XRTX 4090241008
GDDR7RTX 50608 to 16448
GDDR7RTX 507012 to 16672
GDDR7RTX 508016960
GDDR7RTX 5090321792
GDDR7RTX PRO 6000961792

The live table below covers the remaining families, including MI300X. Match the family name to the inventory before comparing prices.

GPUCheapest $/GPU-hrProviderProviders in stock
A30none in stock
Tesla P100none in stock
TITAN Vnone in stock
V100$0.83Ori2
A100$0.68LeaderGPU11
Intel Gaudi 2$0.91LeaderGPU1
MI250Xnone in stock
GH200 Grace Hoppernone in stock
MI300X$3.39Hot Aisle1
B300$7.89RunPod3
GB300none in stock
MI325Xnone in stock
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

Choose GDDR unless capacity or measured latency rules it out

Rent a GDDR card when your full allocation fits and it meets your decode latency and concurrency targets. Pay for HBM when the GDDR candidates fail one of those requirements and the HBM candidate passes. Keep the model, precision, context and request load fixed during that comparison. Choose the least expensive configuration that passes, then increase capacity or bandwidth only for a demonstrated constraint.

Sources

Frequently asked questions

What is HBM memory?▾

HBM is stacked memory beside a processor, connected through a very wide interface. It is used on data-center GPU families such as H100, H200 and B200.

Is GDDR7 enough for LLM inference?▾

Yes, when the model, KV cache and runtime allocations fit and the card meets your latency target. Choose using total memory capacity and bandwidth, then test your intended workload.

Does HBM always generate tokens faster?▾

No. Bandwidth matters most when memory traffic limits decode; larger batches can shift some operations toward a compute limit.

Are HBM3E and HBM4 the same generation?▾

No. Micron specifies a 1024-I/O HBM3E interface, while JEDEC's HBM4 standard specifies 2048 bits per stack.

Does unified memory replace HBM?▾

Unified memory describes a shared system memory pool. Compare its bandwidth and usable allocation separately before treating it as a substitute for dedicated GPU memory.

Related Posts