Data parallelism copies the model; sharded data parallelism, including FSDP and ZeRO, splits its states. Tensor parallelism splits each layer and runs expensive all-reduces, so keep it inside a server on NVLink by default; pipeline parallelism splits by layers and relies on cheaper point-to-point transfers between nodes, while expert parallelism splits mixture-of-experts layers. Combine them by fitting layers within the fastest communication domain first, extending with pipeline stages, then adding data workers for throughput, with expert groups sized for the MoE layout.
That rule follows PyTorch's and DeepSpeed's documentation and NVIDIA's Megatron guidance of 12 April 2021. Start with memory, then draw the communication groups before choosing a multi GPU instance.
Match the split to the traffic
The definitions below come from PyTorch and DeepSpeed documentation retrieved 28 September 2026, Megatron-LM's 17 September 2019 paper, Megatron's 2021 paper revised 23 August 2021, GPipe's 25 July 2019 revision, and SGLang's expert documentation retrieved 28 September 2026. Boundaries are planning recommendations based on those communication patterns, not software restrictions.
| Strategy | What is split | Communication pattern | Typical boundary | Memory saving | Main cost |
|---|---|---|---|---|---|
| Data parallel, DDP | Training examples; model replicated | Gradient all-reduce | Inside or across nodes | No model-state sharding | Gradient synchronization |
| FSDP / ZeRO-3 | Parameters, gradients, optimizer state | Parameter all-gather; gradient reduce-scatter | Inside or across nodes | Shards persistent model states | Repeated parameter movement and transient full parameters |
| ZeRO-1 / ZeRO-2 | Optimizer state / optimizer plus gradients | Partition-dependent reductions and parameter collection | Inside or across nodes | Partial state savings | Remaining replicated states and communication |
| Tensor, TP | Computation and parameters within layers | Frequent layer collectives | Inside an NVLink domain | Partitions layer state | Expensive all-reduces and smaller matrix operations |
| Pipeline, PP | Groups of layers | Neighbor-to-neighbor activations and backward gradients | Across nodes after local capacity is used | Keeps only assigned layers' states | Idle pipeline bubbles and stage imbalance |
| Expert, EP | Experts within MoE layers | Token dispatch and output combination, all-to-all | Inside or across nodes | Distributes expert weights | Routing traffic across the expert group |
Data parallelism: replicate only after the model fits
PyTorch's DDP API, retrieved 28 September 2026, defines the operation as synchronizing gradients across model replicas. Your application must divide the input data, for example with DistributedSampler. DDP does not do that split for you.
PyTorch's design note, retrieved the same day, describes gradient buckets whose asynchronous all-reduces can overlap backward computation. For PyTorch distributed training, this makes DDP a useful starting point when the complete training state already fits. The analyst inference from its replica design is simple: adding replicas does not divide parameter storage. Treat DDP as a throughput choice, not a cure for a model that exceeds device memory.
FSDP and DeepSpeed ZeRO: remove replicated states
PyTorch's FSDP2 API, retrieved 28 September 2026, shards parameters, gradients and optimizer states. It gathers parameters before forward computation and reduce-scatters gradients afterward. With reshard_after_forward=True, it releases full parameters after forward and gathers them again for backward.
DeepSpeed's ZeRO documentation, retrieved 28 September 2026, separates the savings into stages. Stage 1 partitions optimizer state. Stage 2 also partitions gradients. Stage 3 also partitions parameters. Choose the least extensive partitioning that meets your memory budget, then measure whether its communication is acceptable.
DeepSpeed also says ordinary ZeRO-3 execution requires each submodule's forward and backward to fit in device memory. A small average shard does not prove that the largest gathered module fits. This is where tensor splitting can still be necessary.
Tensor parallelism: split an oversized layer
Megatron-LM's 17 September 2019 paper described an intra-layer approach that distributes transformer parameters and computation. Its original implementation uses two all-reduces in forward and two in backward per transformer layer. That count belongs to that implementation, not every modern TP kernel.
NVIDIA's 12 April 2021 explanation keeps TP inside DGX A100 servers to avoid sending those collectives over slower external links. It also warns that higher TP degrees shrink matrix multiplications and can reduce utilization. Choose the smallest group that solves your layer-memory problem. More GPUs in a TP group are not automatically better.
Pipeline parallelism: split depth, then balance work
GPipe's 25 July 2019 revision assigns subsequences of layers to separate accelerators. It divides a mini-batch into micro-batches and accumulates their gradients before a synchronous update. Communication happens between neighboring partitions.
PipeDream's 27 October 2019 paper describes 1F1B as alternating forward and backward passes in steady state. Megatron's 2021 paper, revised 23 August 2021, uses a flush schedule with synchronized updates at batch boundaries. Do not assume every schedule called 1F1B has identical update semantics.
Megatron gives the pipeline bubble relative to ideal compute time as (p - 1) / m, assuming uniformly timed stages, with p stages and m micro-batches. Derived: four stages and sixteen micro-batches give 3 / 16 = 18.75%. PipeDream identifies the slowest stage as the throughput bottleneck. Balance stage compute before adding stages just to lower memory per GPU.
Expert parallelism: route tokens to distributed experts
NVIDIA's Megatron guide, retrieved 28 September 2026, defines EP as distributing MoE experts across GPUs. SGLang's expert documentation, retrieved the same day, describes all-to-all token dispatch and output combination. The relevant rental boundary is therefore the whole expert group, including its cross-node paths.
NVIDIA's guide also requires sequence parallelism when combining TP and EP in Megatron. Use the framework's supported layout; a plausible diagram alone is not a valid configuration.
How many GPUs: a derived 7B training budget
Use an illustrative dense model with 7 billion parameters. The ZeRO paper's 13 May 2020 accounting assumes FP16 parameters and gradients plus FP32 Adam master weights, momentum and variance. It totals 16 bytes per parameter, giving a derived 112 decimal GB of unsharded model states: 7 billion × 16.
For this planning example, choose H100 SXM with 80 GB from NVIDIA's specification as recorded in the site's verified spec data. Assign an illustrative 60 GB budget to model states on each GPU. The remaining capacity is a planning allowance, not a measured requirement or a promise that runtime allocations fit.
The following figures are derived from ZeRO's formulas. Here n is the number of workers. TP and PP rows assume perfectly balanced partitioning of all model states, without replicated exceptions; EP does not apply to this dense example.
| Strategy alone | Model-state GB per GPU | Smallest count within the assumed 60 GB budget |
|---|---|---|
| DDP | 112, regardless of replicas | No count solves the per-GPU limit |
| ZeRO-1 | 28 + 84/n | 3 GPUs: 56 GB; 2 would need 70 GB |
| ZeRO-2 | 14 + 98/n | 3 GPUs: 46.7 GB; 2 would need 63 GB |
| ZeRO-3 / ideal full FSDP sharding | 112/n | 2 GPUs: 56 GB persistent states |
| Ideal TP state partition | 112/n | 2 GPUs: 56 GB, subject to supported layer splits |
| Ideal PP state partition | 112/n | 2 GPUs: 56 GB, subject to balanced stages |
| EP | Not applicable to a dense model | No expert-based reduction |
ZeRO explicitly treats activations, temporary buffers and fragmented memory as additional requirements. These counts are arithmetic starting points, not verified minimum rentals. FSDP's gathered parameters need space too. Test the intended sequence length and micro-batch before committing to a longer run.
For multi GPU training, start by testing full sharding at the derived count, then increase capacity if peak memory demands it. Use the LLM VRAM calculator for a separate serving estimate. If your task permits adapter training, compare LoRA and QLoRA before budgeting full training states.
Published layouts show how the groups combine
Megatron's 2021 paper, revised 23 August 2021, reports a 1,008-billion-parameter configuration on 3,072 GPUs, with TP=8 and PP=64. Derived: 3,072 / (8 × 64) = 6 data workers. This is a published scaling experiment, not evidence of completed pretraining or a fixed-model speedup.
Meta's Llama 3 paper, revised 23 November 2024, calls its combination "4D parallelism": tensor, context, pipeline and fully sharded data parallelism. Its 405B table includes 16,384 GPUs with TP=8, CP=1, PP=16 and DP=128. The long-context row keeps that GPU count but uses CP=16 and DP=8, showing how sequence length changes the layout.
DeepSeek's 27 December 2024 V3 report lists 2,048 H800 GPUs, "16-way Pipeline Parallelism", "64-way Expert Parallelism" spanning eight nodes, and ZeRO-1. It reports memory optimizations that let training omit TP. Do not transfer that conclusion to serving: DeepSeek's reported prefill deployment uses attention TP4, SP and DP8, with EP32 across 32 GPUs.
Serving: set sizes in the engine's terms
vLLM's scaling documentation, retrieved 28 September 2026, recommends one GPU when the model fits and TP when it fits one node. For a model that exceeds a node, its example uses TP=8 and PP=2 across two eight-GPU nodes. vLLM also documents cross-node TP=16, so the local-TP rule is a default, not a hard limit.
vLLM selects TP with --tensor-parallel-size. Its startup KV-cache capacity and concurrency estimates help you judge fit at your intended request lengths. Weight storage alone is not the serving capacity target.
vLLM's expert deployment guide, retrieved 28 September 2026, enables EP with --enable-expert-parallel and calculates EP size as TP×DP. Its TP=2, DP=4 example uses eight GPUs in one expert group. Do not multiply by EP again; that would count the same ranks twice.
SGLang's argument documentation, retrieved 28 September 2026, exposes --tensor-parallel-size and --pipeline-parallel-size. Its pipeline guide describes asynchronous point-to-point transfers at stage boundaries and gives a four-node TP=8, PP=4 long-context example. Its expert guide gives a separate DeepSeek-V3 example using --tp 8 --ep 8, with DeepEP communication and deep_gemm computation. These are documented layouts, not universal minimum GPU counts. Verify flags against your installed release; the serving-engine comparison helps narrow the implementation choice.
Rent the communication domain, then the GPUs
Compare H100, H200, B200 and A100 capacity and links here:
Before renting, ask for the actual NVLink group and the network path between groups. Use the NVLink, PCIe and SXM guide to interpret the listing. For rack-scale proposals, check the domain in the GB200 and GB300 NVL72 guide rather than assuming a chassis defines it.
Meta's 12 March 2024 infrastructure post describes both RoCE and InfiniBand clusters. Meta's own conclusion was that both supported large GenAI workloads without network bottlenecks. That supports asking about the implemented fabric, not choosing solely by its name. Use the InfiniBand and RoCE comparison to frame that discussion.
The live table below shows SXM rental listings. Request multi-node cluster quotes with your TP group, pipeline layout and network requirements attached.
The decision rule: use DDP when training states fit; shard states when they do not; use TP for oversized layers inside the fastest domain; extend with balanced pipeline stages across nodes. For MoE, size and price the expert communication group explicitly. Rent the smallest layout that passes your memory and throughput checks.
Sources
- PyTorch DDP API, retrieved 28 September 2026.
- PyTorch DDP design, retrieved 28 September 2026.
- PyTorch FSDP2, retrieved 28 September 2026.
- DeepSpeed ZeRO, retrieved 28 September 2026.
- ZeRO paper, 13 May 2020.
- Megatron-LM original paper, 17 September 2019.
- NVIDIA Megatron scaling guidance, 12 April 2021.
- Megatron distributed training paper, revision 23 August 2021.
- GPipe, revision 25 July 2019.
- PipeDream, 27 October 2019.
- NVIDIA Megatron parallelism guide, retrieved 28 September 2026.
- Meta Llama 3 paper, revision 23 November 2024.
- DeepSeek-V3 technical report, 27 December 2024.
- vLLM parallelism and scaling, retrieved 28 September 2026.
- vLLM expert deployment, retrieved 28 September 2026.
- SGLang server arguments, retrieved 28 September 2026.
- SGLang pipeline parallelism, retrieved 28 September 2026.
- SGLang expert parallelism, retrieved 28 September 2026.
- Meta GenAI infrastructure, 12 March 2024.