8 articles in this category
Pick a precision for pre-training, fine-tuning or inference, and see which rentable GPUs accelerate FP8 and FP4 in hardware and which only emulate them.
VRAM, memory bandwidth and dense throughput decide an AI workload. How to find them on a datasheet, and why the headline TFLOPS figure is usually doubled.
A is Ampere, L is Ada, H is Hopper, B is Blackwell. A decoder table, a generation timeline and a four-step method for placing any NVIDIA GPU on a rental site.
Verified purchase listings for NVIDIA H200, B200, B300, GB200 NVL72 and A100 with seller and date, the misquoted figures corrected, and live rental prices.
What an H100 card, an 8-GPU server and a DGX H100 last listed for, with sellers and dates, next to live rental prices and a break-even method you can run.
NVLink, PCIe and SXM explained for people renting multi-GPU servers: bandwidth per generation, what the H100 NVL is, and when the SXM premium pays off.
What a Google TPU is, what each generation costs per chip-hour in September 2026, how that sits beside live GPU rental prices, and when to pick which.
MIG splits one NVIDIA data centre GPU into isolated instances with their own memory. See which GPUs support it and when a slice beats a whole cheap GPU.