AWS GPU Instances Decoded Across Four Major Clouds

Decode AWS, Google Cloud, Azure and Oracle GPU names, check GPU counts and memory, then compare live AWS rates and dated Oracle rates with other providers.

By Faiz Ahmed•
•13 min read

AWS encodes series and size, Google encodes machine series and configuration, Azure encodes family, vCPUs and sometimes the accelerator, and Oracle uses GPU shape names, according to their naming documentation listed below. For AI, start with AWS P and G, Google A and G, Azure ND and NC, and Oracle BM.GPU. Uptime Institute wrote on 26 February 2025 that neoclouds, which focus on GPU-backed servers and virtual machines, often charge less than the hyperscalers; use the live tables below to compare today's offers.

Decode the machine before comparing its price. Write down the exact GPU, GPU count and memory per GPU. Then check the whole-machine bill. A low per-GPU figure is useful only if you can use the allocation you must rent.

These tables focus on H100, H200, A100, L4 and A10-class comparisons. Memory means nominal capacity per GPU, from the site's verified GPU specification reference, not host RAM or a promised usable VM allocation. Cloud names and GPU counts come from the publishers identified beside each table. Only the AWS table includes network figures; leaving them out of the other tables is not a claim of zero bandwidth.

AWS GPU instances: decode P and G before the size

AWS's naming guide, accessed 28 September 2026, puts the series, generation and optional capabilities before the period. The size follows it. AWS defines P as GPU accelerated and G as graphics intensive. Use the family mapping to identify the GPU; the generation digit alone does not name the NVIDIA product.

AWS name fieldMeaning in AWS's naming guideReading example
p or gGPU accelerated or graphics intensiveStart with the family
Generation digitGeneration within the seriesp5 is a family identifier
eExtra GPU memoryPresent in p5e
nNetwork and EBS optimizedPresent in p5en
dInstance-store volumesA capability suffix
.48xlargeInstance sizeNot 48 GPUs
.metalBare-metal sizeDeployment form

AWS's accelerated-computing reference, accessed 28 September 2026, gives the following GPU allocations and P-family network figures. AWS's G5 and G6 product pages, accessed the same day, supply their network figures. GPU memory below uses nominal GB from the verified specification reference, without converting AWS's GiB allocation labels.

EC2 instanceGPUGPU countNominal memory per GPUAWS network figure
p4d.24xlargeA100840 GB4 × 100 Gigabit
p4de.24xlargeA100880 GB4 × 100 Gigabit
p5.4xlargeH100180 GB100 Gigabit
p5.48xlargeH100880 GB3200 Gigabit
p5e.48xlargeH2008141 GB3200 Gigabit
p5en.48xlargeH2008141 GB3200 Gigabit
g5.xlargeA10G1Confirm A10G allocationUp to 10 Gbps
g5.48xlargeA10G8Confirm A10G allocation100 Gbps
g6.xlargeL4124 GBUp to 10 Gbps
g6.12xlargeL4424 GB40 Gbps
g6.48xlargeL4824 GB100 Gbps

For an AWS H100 instance, AWS's reference makes the size distinction concrete: p5.4xlarge has one GPU and p5.48xlarge has eight. Do not choose the larger AWS P5 instance just because its normalized rate looks attractive. First decide whether your job needs the whole allocation.

Keep the suffix when copying EC2 GPU instances into a budget. AWS maps P5e and P5en to H200, while plain P5 maps to H100. For G5, request the exact A10G memory allocation before booking. The live A10 comparison is a family-level price reference, not proof that an A10 offer reproduces an A10G instance.

For a network label marked "up to," do not budget around sustained maximum throughput. AWS's reference, accessed 28 September 2026, identifies credit-based bursting on those sizes. Confirm the baseline for the exact instance. Keep the AWS provider page beside your quote so you can inspect the tracked offer rather than treating a family minimum as the price of every size.

Google Cloud GPU names: read the entire machine type

Google's machine-resource guide, accessed 28 September 2026, distinguishes family, series and machine type, and says a machine type specifies the instance's resource configuration. This matters because similarly shaped numeric suffixes describe different resources.

Google name field or exampleDocumented interpretation
Family, series, machine typeThree separate terms; a machine type specifies the instance's resource configuration
a3-highgpu-8gGoogle's table maps this type to eight GPUs
g2-standard-32Google's table maps this type to 32 vCPUs and one GPU
-metalMachine type without a hypervisor

Google's accelerator-optimized machine documentation, dated 28 September 2026, gives these mappings. The rows deliberately range from one GPU to eight so you can distinguish the GPU from the amount you rent.

Google machine typeGPUGPU countNominal memory per GPU
a2-highgpu-1gA100140 GB
a2-highgpu-8gA100840 GB
a2-ultragpu-1gA100180 GB
a2-ultragpu-8gA100880 GB
a3-highgpu-1gH100 SXM180 GB
a3-highgpu-8gH100 SXM880 GB
a3-megagpu-8gH100 SXM880 GB
a3-ultragpu-8gH200 SXM8141 GB
g2-standard-4L4124 GB
g2-standard-24L4224 GB
g2-standard-32L4124 GB
g2-standard-96L4824 GB

For a GCP GPU quote, copy the middle part of the name too. Google's table assigns H100 to A3 High and Mega, but H200 to A3 Ultra. Treating all A3 machines as interchangeable loses the GPU identity before the price comparison even starts.

For Google Cloud GPU pricing, request the complete machine charge. Google's documentation dated 28 September 2026 says accelerator-optimized billing includes attached GPUs, predefined vCPUs, host memory and bundled Local SSD where applicable. Divide that quote by its GPU count before comparing it with a per-GPU listing. Then check what the competing listing includes. No Google list price is reproduced here.

Azure GPU VM names: the large number means vCPUs

Microsoft's naming convention, dated 1 December 2025, identifies the numeric field as vCPUs, not GPUs. It defines ND as AI training and inference optimized and NC as compute intensive. The accelerator field can name the hardware, but older names still require a lookup.

Azure name fieldMicrosoft definition or reading
ND, NCAI training/inference optimized; compute intensive
96 in ND96vCPU count
aAMD CPU
dLocal temporary disks
mMemory-intensive variant
rRDMA/InfiniBand secondary network
sPremium SSD compatibility
H100Accelerator identifier
v5VM-series version

Microsoft's size documentation dated 28 July 2026 supplies these allocations, except the ND H100 page, dated 25 September 2026. Do not infer a numeric network rate from the presence of r; it identifies a feature, not a speed.

Azure sizeGPUGPU countNominal memory per GPU
Standard_ND96asr_v4A100840 GB
Standard_ND96amsr_A100_v4A100880 GB
Standard_ND96isr_H100_v5H100880 GB
Standard_ND96isr_H200_v5H2008141 GB
Standard_NC24ads_A100_v4A100 PCIe180 GB
Standard_NC48ads_A100_v4A100 PCIe280 GB
Standard_NC96ads_A100_v4A100 PCIe480 GB
Standard_NC40ads_H100_v5H100 NVL194 GB
Standard_NC80adis_H100_v5H100 NVL294 GB

An H100 label is not enough to match two offers. Microsoft's ND and NC rows above specify different H100 configurations. Use the NVLink, PCIe and SXM guide when checking whether a cheaper listing matches the topology your workload requires.

For Azure GPU pricing, obtain a quote for the exact size, region and billing term. Keep it separate from any account discount until both are visible in your comparison. This article gives no numeric Azure rate and makes no price ranking between Azure and Google.

Oracle BM.GPU names: distinguish the shape from the bill

Oracle's shape reference, accessed 28 September 2026, separates bare-metal and VM GPU shapes. Its explicit mapping makes BM.GPU.H100.8 easy to decode, but do not generalize that example into a rule for every trailing number in OCI.

Oracle name fieldMeaning in Oracle's shape reference
BM.GPUBare-metal GPU shape category
VM.GPUVirtual-machine GPU shape category
H100 in BM.GPU.H100.8GPU model
.8 in BM.GPU.H100.8Eight GPUs in this documented shape

Oracle's same reference gives the allocations below. Use the full shape name when asking for a quote, including the A100 version marker.

Oracle shapeGPUGPU countNominal memory per GPU
BM.GPU.H100.8H100880 GB
BM.GPU.H200.8H2008141 GB
BM.GPU.A100-v2.8A100880 GB
VM.GPU.A10.1A10124 GB
VM.GPU.A10.2A10224 GB
BM.GPU.A10.4A10424 GB

Oracle's global price list dated 10 September 2026, observed 28 September 2026 with country USA and currency USD selected, publishes these pay-as-you-go compute rates. The geographic scope is the USA-selected global schedule, not a quote for a named OCI region.

Oracle shapePublished USD per GPU-hourDate and geographic scope
BM.GPU.H100.8USD 10.0010 September 2026; USA-selected global list
BM.GPU.H200.8USD 10.0010 September 2026; USA-selected global list
VM.GPU.A10.1, VM.GPU.A10.2, BM.GPU.A10.4USD 2.0010 September 2026; USA-selected global list

Oracle's price list, accessed 28 September 2026, defines the whole-server hourly charge as the listed GPU rate multiplied by the server's GPU count. Compare the published per-GPU rate with the matching live row below, then price the allocation you actually need. These dated list rates do not confirm capacity in your chosen region.

Compare live AWS GPU pricing with the same GPU elsewhere

Read across the AWS row, then find that GPU in the cheapest-offer table. Both express prices per GPU. The difference is your starting point for comparing suppliers, not a guaranteed saving for a particular instance or job.

ProviderH100 $/GPU-hrH200 $/GPU-hrA100 $/GPU-hrL4 $/GPU-hrA10 $/GPU-hr
AWSnone in stocknone in stocknone in stocknone in stocknone in stock
Cheapest in-stock on-demand price per GPU-hour at each provider, from providers with live stock tracking. Latest stock observation: .
GPUCheapest $/GPU-hrProviderProviders in stock
H100$2.50Hyperstack9
H200$3.43QuantaCloud6
A100$0.68LeaderGPU11
L4$0.49RunPod2
A10$0.37LeaderGPU2
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Build a short quote worksheet with the exact machine name, region, GPU variant, allocation, billing term and expected runtime. Keep the per-GPU comparison beside the whole-job estimate. If a listing leaves a field unclear, resolve it before selecting the winner. This also gives your team a repeatable way to revisit the decision when either the workload or the offer changes.

Match memory variant, GPU count, rental term and included host resources before using the difference. For H100, inspect the variants directly:

VariantVRAMCheapest $/GPU-hrProviderProviders in stock
H100 PCIe80 GB$2.50Hyperstack6
H100 SXM580 GB$2.90Ori4
H100 NVL94 GB$3.11Massed Compute3
Cheapest in-stock on-demand price per GPU-hour for each H100 variant, from providers with live stock tracking. Latest stock observation: .

If you have an AWS contract quote, compare that quote too. A public offer does not tell you the value of your credits or negotiated terms. Use the on-demand, spot and reserved pricing guide to keep different commitments out of the same comparison column.

Check provisioning before choosing a supplier

AWS's Capacity Blocks documentation, accessed 21 September 2026, describes reserving GPU instances for a future date and says cancellations are not allowed. Confirm the start and end of the block against your job plan before committing.

Google's documentation dated 28 September 2026 requires Spot or Flex-start provisioning for a3-highgpu-1g, a3-highgpu-2g and a3-highgpu-4g. The same documentation limits accelerator types to particular regions and zones. A machine name therefore does not establish that your preferred provisioning mode exists in your region.

Microsoft's RTX PRO 6000 BSE v6 overview dated 1 September 2026 reports general availability in West US 2 and Southeast Asia. Treat that as a dated regional announcement, not evidence of spare capacity today. For every cloud, make region, allocation and start time part of the quote request.

Keep the cloud only when its total cost wins

Stay inside your hyperscaler when usable credits, data gravity, compliance requirements or existing contracts outweigh the compute-price difference. Put a value on each reason. Include transfer and storage in the calculation using the data-egress reference. Count the engineering work required to move the job and operate it elsewhere.

Rent the same GPU elsewhere when the workload is portable, the offer meets your requirements and the complete job costs less. The neocloud guide explains that provider category. If your proposed move also changes the accelerator, use the AWS Trainium versus GPU comparison; that is a separate decision from changing who rents you the GPU.

Choose the smallest matching allocation, confirm its terms, and compare total job cost. Stay for a concrete financial or operational advantage. Otherwise, rent the matching GPU from the cheaper qualified offer.

Sources

Frequently asked questions

How many GPUs does p5.48xlarge contain?▾

AWS lists eight H100 GPUs for p5.48xlarge. The 48xlarge suffix is an instance size, not a GPU count.

Does every Google A3 machine use H100 GPUs?▾

No. Google lists H100 SXM for A3 High and A3 Mega, but H200 SXM for A3 Ultra. Check the complete machine type.

Does 96 in an Azure ND name mean 96 GPUs?▾

No. Microsoft's naming convention uses that field for vCPUs. Standard_ND96isr_H100_v5 contains eight H100 GPUs.

Are Oracle's published rates for an entire GPU server?▾

The Oracle rates shown here are per GPU. Oracle's price list defines the whole-server hourly charge as that rate multiplied by the server's GPU count.

Should I move a GPU job outside my hyperscaler?▾

Move when the complete job is cheaper after transfer, storage and operating work, and the provider meets your requirements. Stay when credits, data location, compliance or an existing contract make your current cloud the better total-cost choice.

Related Posts