Tensor vs Pipeline vs Data Parallelism: How to Split a Model

Choose how to split a model across rented GPUs, calculate training memory, and match tensor, pipeline and data parallelism to NVLink and cluster networks.

By Faiz Ahmed•
•11 min read

Data parallelism copies the model; sharded data parallelism, including FSDP and ZeRO, splits its states. Tensor parallelism splits each layer and runs expensive all-reduces, so keep it inside a server on NVLink by default; pipeline parallelism splits by layers and relies on cheaper point-to-point transfers between nodes, while expert parallelism splits mixture-of-experts layers. Combine them by fitting layers within the fastest communication domain first, extending with pipeline stages, then adding data workers for throughput, with expert groups sized for the MoE layout.

That rule follows PyTorch's and DeepSpeed's documentation and NVIDIA's Megatron guidance of 12 April 2021. Start with memory, then draw the communication groups before choosing a multi GPU instance.

Match the split to the traffic

The definitions below come from PyTorch and DeepSpeed documentation retrieved 28 September 2026, Megatron-LM's 17 September 2019 paper, Megatron's 2021 paper revised 23 August 2021, GPipe's 25 July 2019 revision, and SGLang's expert documentation retrieved 28 September 2026. Boundaries are planning recommendations based on those communication patterns, not software restrictions.

StrategyWhat is splitCommunication patternTypical boundaryMemory savingMain cost
Data parallel, DDPTraining examples; model replicatedGradient all-reduceInside or across nodesNo model-state shardingGradient synchronization
FSDP / ZeRO-3Parameters, gradients, optimizer stateParameter all-gather; gradient reduce-scatterInside or across nodesShards persistent model statesRepeated parameter movement and transient full parameters
ZeRO-1 / ZeRO-2Optimizer state / optimizer plus gradientsPartition-dependent reductions and parameter collectionInside or across nodesPartial state savingsRemaining replicated states and communication
Tensor, TPComputation and parameters within layersFrequent layer collectivesInside an NVLink domainPartitions layer stateExpensive all-reduces and smaller matrix operations
Pipeline, PPGroups of layersNeighbor-to-neighbor activations and backward gradientsAcross nodes after local capacity is usedKeeps only assigned layers' statesIdle pipeline bubbles and stage imbalance
Expert, EPExperts within MoE layersToken dispatch and output combination, all-to-allInside or across nodesDistributes expert weightsRouting traffic across the expert group

Data parallelism: replicate only after the model fits

PyTorch's DDP API, retrieved 28 September 2026, defines the operation as synchronizing gradients across model replicas. Your application must divide the input data, for example with DistributedSampler. DDP does not do that split for you.

PyTorch's design note, retrieved the same day, describes gradient buckets whose asynchronous all-reduces can overlap backward computation. For PyTorch distributed training, this makes DDP a useful starting point when the complete training state already fits. The analyst inference from its replica design is simple: adding replicas does not divide parameter storage. Treat DDP as a throughput choice, not a cure for a model that exceeds device memory.

FSDP and DeepSpeed ZeRO: remove replicated states

PyTorch's FSDP2 API, retrieved 28 September 2026, shards parameters, gradients and optimizer states. It gathers parameters before forward computation and reduce-scatters gradients afterward. With reshard_after_forward=True, it releases full parameters after forward and gathers them again for backward.

DeepSpeed's ZeRO documentation, retrieved 28 September 2026, separates the savings into stages. Stage 1 partitions optimizer state. Stage 2 also partitions gradients. Stage 3 also partitions parameters. Choose the least extensive partitioning that meets your memory budget, then measure whether its communication is acceptable.

DeepSpeed also says ordinary ZeRO-3 execution requires each submodule's forward and backward to fit in device memory. A small average shard does not prove that the largest gathered module fits. This is where tensor splitting can still be necessary.

Tensor parallelism: split an oversized layer

Megatron-LM's 17 September 2019 paper described an intra-layer approach that distributes transformer parameters and computation. Its original implementation uses two all-reduces in forward and two in backward per transformer layer. That count belongs to that implementation, not every modern TP kernel.

NVIDIA's 12 April 2021 explanation keeps TP inside DGX A100 servers to avoid sending those collectives over slower external links. It also warns that higher TP degrees shrink matrix multiplications and can reduce utilization. Choose the smallest group that solves your layer-memory problem. More GPUs in a TP group are not automatically better.

Pipeline parallelism: split depth, then balance work

GPipe's 25 July 2019 revision assigns subsequences of layers to separate accelerators. It divides a mini-batch into micro-batches and accumulates their gradients before a synchronous update. Communication happens between neighboring partitions.

PipeDream's 27 October 2019 paper describes 1F1B as alternating forward and backward passes in steady state. Megatron's 2021 paper, revised 23 August 2021, uses a flush schedule with synchronized updates at batch boundaries. Do not assume every schedule called 1F1B has identical update semantics.

Megatron gives the pipeline bubble relative to ideal compute time as (p - 1) / m, assuming uniformly timed stages, with p stages and m micro-batches. Derived: four stages and sixteen micro-batches give 3 / 16 = 18.75%. PipeDream identifies the slowest stage as the throughput bottleneck. Balance stage compute before adding stages just to lower memory per GPU.

Expert parallelism: route tokens to distributed experts

NVIDIA's Megatron guide, retrieved 28 September 2026, defines EP as distributing MoE experts across GPUs. SGLang's expert documentation, retrieved the same day, describes all-to-all token dispatch and output combination. The relevant rental boundary is therefore the whole expert group, including its cross-node paths.

NVIDIA's guide also requires sequence parallelism when combining TP and EP in Megatron. Use the framework's supported layout; a plausible diagram alone is not a valid configuration.

How many GPUs: a derived 7B training budget

Use an illustrative dense model with 7 billion parameters. The ZeRO paper's 13 May 2020 accounting assumes FP16 parameters and gradients plus FP32 Adam master weights, momentum and variance. It totals 16 bytes per parameter, giving a derived 112 decimal GB of unsharded model states: 7 billion × 16.

For this planning example, choose H100 SXM with 80 GB from NVIDIA's specification as recorded in the site's verified spec data. Assign an illustrative 60 GB budget to model states on each GPU. The remaining capacity is a planning allowance, not a measured requirement or a promise that runtime allocations fit.

The following figures are derived from ZeRO's formulas. Here n is the number of workers. TP and PP rows assume perfectly balanced partitioning of all model states, without replicated exceptions; EP does not apply to this dense example.

Strategy aloneModel-state GB per GPUSmallest count within the assumed 60 GB budget
DDP112, regardless of replicasNo count solves the per-GPU limit
ZeRO-128 + 84/n3 GPUs: 56 GB; 2 would need 70 GB
ZeRO-214 + 98/n3 GPUs: 46.7 GB; 2 would need 63 GB
ZeRO-3 / ideal full FSDP sharding112/n2 GPUs: 56 GB persistent states
Ideal TP state partition112/n2 GPUs: 56 GB, subject to supported layer splits
Ideal PP state partition112/n2 GPUs: 56 GB, subject to balanced stages
EPNot applicable to a dense modelNo expert-based reduction

ZeRO explicitly treats activations, temporary buffers and fragmented memory as additional requirements. These counts are arithmetic starting points, not verified minimum rentals. FSDP's gathered parameters need space too. Test the intended sequence length and micro-batch before committing to a longer run.

For multi GPU training, start by testing full sharding at the derived count, then increase capacity if peak memory demands it. Use the LLM VRAM calculator for a separate serving estimate. If your task permits adapter training, compare LoRA and QLoRA before budgeting full training states.

Published layouts show how the groups combine

Megatron's 2021 paper, revised 23 August 2021, reports a 1,008-billion-parameter configuration on 3,072 GPUs, with TP=8 and PP=64. Derived: 3,072 / (8 × 64) = 6 data workers. This is a published scaling experiment, not evidence of completed pretraining or a fixed-model speedup.

Meta's Llama 3 paper, revised 23 November 2024, calls its combination "4D parallelism": tensor, context, pipeline and fully sharded data parallelism. Its 405B table includes 16,384 GPUs with TP=8, CP=1, PP=16 and DP=128. The long-context row keeps that GPU count but uses CP=16 and DP=8, showing how sequence length changes the layout.

DeepSeek's 27 December 2024 V3 report lists 2,048 H800 GPUs, "16-way Pipeline Parallelism", "64-way Expert Parallelism" spanning eight nodes, and ZeRO-1. It reports memory optimizations that let training omit TP. Do not transfer that conclusion to serving: DeepSeek's reported prefill deployment uses attention TP4, SP and DP8, with EP32 across 32 GPUs.

Serving: set sizes in the engine's terms

vLLM's scaling documentation, retrieved 28 September 2026, recommends one GPU when the model fits and TP when it fits one node. For a model that exceeds a node, its example uses TP=8 and PP=2 across two eight-GPU nodes. vLLM also documents cross-node TP=16, so the local-TP rule is a default, not a hard limit.

vLLM selects TP with --tensor-parallel-size. Its startup KV-cache capacity and concurrency estimates help you judge fit at your intended request lengths. Weight storage alone is not the serving capacity target.

vLLM's expert deployment guide, retrieved 28 September 2026, enables EP with --enable-expert-parallel and calculates EP size as TP×DP. Its TP=2, DP=4 example uses eight GPUs in one expert group. Do not multiply by EP again; that would count the same ranks twice.

SGLang's argument documentation, retrieved 28 September 2026, exposes --tensor-parallel-size and --pipeline-parallel-size. Its pipeline guide describes asynchronous point-to-point transfers at stage boundaries and gives a four-node TP=8, PP=4 long-context example. Its expert guide gives a separate DeepSeek-V3 example using --tp 8 --ep 8, with DeepEP communication and deep_gemm computation. These are documented layouts, not universal minimum GPU counts. Verify flags against your installed release; the serving-engine comparison helps narrow the implementation choice.

Rent the communication domain, then the GPUs

Compare H100, H200, B200 and A100 capacity and links here:

SpecH100H200B200A100
VRAM80 to 94 GB141 GB180 to 192 GB40 to 80 GB
Memory bandwidth3,350 GB/s4,800 GB/s8,000 GB/s2,039 GB/s
InterconnectNVLink, PCIe 5.0, InfiniBandNVLink, PCIe 5.0, InfiniBandNVLink, PCIe 6.0, InfiniBandNVLink, PCIe 4.0, InfiniBand
Figures from the vendor datasheets: H100, H200, B200, A100, checked 13 Sep 2026. "Not published" means the vendor gives no figure.

Before renting, ask for the actual NVLink group and the network path between groups. Use the NVLink, PCIe and SXM guide to interpret the listing. For rack-scale proposals, check the domain in the GB200 and GB300 NVL72 guide rather than assuming a chassis defines it.

Meta's 12 March 2024 infrastructure post describes both RoCE and InfiniBand clusters. Meta's own conclusion was that both supported large GenAI workloads without network bottlenecks. That supports asking about the implemented fabric, not choosing solely by its name. Use the InfiniBand and RoCE comparison to frame that discussion.

The live table below shows SXM rental listings. Request multi-node cluster quotes with your TP group, pipeline layout and network requirements attached.

GPUCheapest $/GPU-hrProviderProviders in stock
H100 SXM5$2.90Ori4
H200 SXM$3.50Ori3
B200 SXM$6.79RunPod3
A100 SXM4 80GB$0.80Vast.ai6
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

The decision rule: use DDP when training states fit; shard states when they do not; use TP for oversized layers inside the fastest domain; extend with balanced pipeline stages across nodes. For MoE, size and price the expert communication group explicitly. Rent the smallest layout that passes your memory and throughput checks.

Sources

Frequently asked questions

What is the difference between tensor and pipeline parallelism?▾

Tensor parallelism splits computation inside each layer. Pipeline parallelism assigns groups of layers to different GPUs, communicating at stage boundaries.

Does data parallelism reduce model memory per GPU?▾

Ordinary DDP keeps a model replica on each worker. FSDP and ZeRO partition model states to reduce that storage, with additional communication.

How many GPUs do I need for a 7B training job?▾

The derived example needs two GPUs for ideally balanced full state sharding, tensor splitting or pipeline splitting within its assumed memory budget. That is a model-state estimate, not a guarantee that activations and temporary allocations fit.

Must tensor parallelism stay inside one server?▾

No, but it is the default rental recommendation because frequent collectives favor NVLink-class bandwidth. vLLM also documents cross-node tensor parallelism.

Do I multiply expert size by tensor and data sizes?▾

Not in vLLM's documented expert layout: its expert group already spans the tensor and data ranks. Use the framework's group definitions before counting GPUs.

Related Posts