2 articles with this tag
Pick a precision for pre-training, fine-tuning or inference, and see which rentable GPUs accelerate FP8 and FP4 in hardware and which only emulate them.
VRAM, memory bandwidth and dense throughput decide an AI workload. How to find them on a datasheet, and why the headline TFLOPS figure is usually doubled.