2 articles with this tag
Choose how to split a model across rented GPUs, calculate training memory, and match tensor, pipeline and data parallelism to NVLink and cluster networks.
InfiniBand and RoCE Ethernet both move GPU traffic between servers. See the generations, provider fabrics, and when a single GPU or node needs neither.