For ComfyUI cloud use, image models such as SDXL fit modest cards: Comfy Org's guide, as of 28 September 2026, lists 8GB minimum for SDXL Base FP16. Full 16-bit FLUX.1 needs roughly 24 GB for its diffusion weights alone, derived from the parameter count on Black Forest Labs' model card, while Wan's official A14B video commands specify at least 80GB, so choose by the biggest model and workflow you will load. The cheapest suitable GPUs can then be selected from the live table below using those memory requirements.
ComfyUI GPU requirements start with the workflow
Write down the checkpoint, precision, output dimensions, batch size and video length before opening a rental listing. Treat each figure below as applying only to the named path. A smaller image or an offloaded encoder is a different configuration, not evidence that every graph fits the same card.
Dates in the table identify the publisher's documentation as accessed on that date unless a publication date is stated. "Not specified" means you should not infer a missing condition. The table deliberately distinguishes reported runtime memory from weight-only arithmetic.
| Model | Variant or precision | VRAM guidance and limits | Resolution or frames stated | Source and date |
|---|---|---|---|---|
| SDXL | Base FP16 | 8GB minimum; 12GB recommended; batch and offloading unspecified | Native 1024x1024, not a complete memory test setup | Comfy Org model guide, 28 Sep 2026 |
| SDXL | Base plus Refiner | 16GB+ guidance; sequential offloading may change fit | Not specified for this memory figure | Comfy Org model guide, 28 Sep 2026 |
| SD 3.5 Medium | Precision unspecified | VENDOR CLAIM: 9.9 GB excluding text encoders | Model supports 0.25 to 2 megapixels; no memory-test resolution given | Stability AI announcement, 22 Oct 2024, Medium update 29 Oct 2024 |
| SD 3.5 Large | FP16 | 16 to 24GB; offloading unspecified | Not specified | Comfy Org model guide, 28 Sep 2026 |
| FLUX.1 dev and schnell | BF16 diffusion weights only | Derived: about 24 GB, excluding encoders, VAE and runtime allocations | dev model-card example: 1024x1024 with CPU offloading; not a peak-memory measurement | Black Forest Labs model cards, 28 Sep 2026 |
| FLUX.1 dev and schnell | FP8 diffusion weights only | Derived: about 12 GB before metadata and runtime allocations; official examples warn of degraded quality | No complete-workflow minimum established | Black Forest Labs and ComfyUI examples, 28 Sep 2026 |
| HiDream-I1 | Full, Dev or Fast, FP8 / 16-bit | More than 16GB / more than 27GB; FP8 quality equivalence is not established by this fit guidance | Not specified | Comfy Org HiDream guide, 28 Sep 2026 |
| Wan 2.1 | T2V-1.3B | VENDOR CLAIM: 8.19 GB; measurement setup incomplete | Separate official example: 832x480 with model offload and CPU T5 | Wan-Video repository, 28 Sep 2026 |
| Wan 2.1 | 14B, precision unspecified | 24GB+ broad guidance; not a verified production minimum | Frame count and offloading unspecified | Comfy Org Wan overview, 28 Sep 2026 |
| Wan 2.2 | TI2V-5B official CLI | At least 24GB with model offload, dtype conversion and CPU T5 | 1280x704 | Wan-Video repository, 28 Sep 2026 |
| Wan 2.2 | TI2V-5B native ComfyUI FP16 | VENDOR CLAIM: should fit on 8GB with native offloading; not an all-resident fit | No frame-count benchmark attached | Comfy Org Wan 2.2 guide, 28 Sep 2026 |
| Wan 2.2 | T2V-A14B / I2V-A14B official CLI | At least 80GB; T2V command uses offloading and dtype conversion | 1280x720 | Wan-Video repository, 28 Sep 2026 |
| Original HunyuanVideo | Precision unspecified in requirements table | VENDOR CLAIM: 45GB / 60GB peak, batch size 1 | 544x960 / 720x1280, both 129 frames | Tencent Hunyuan repository, 28 Sep 2026 |
| LTX-Video | Original 2B / 13B; FP8 checkpoints also published | No stock-workflow minimum established here; FP8 quality equivalence unestablished | Lightricks' 15 Apr 2025 default: 1216x704 at 30 FPS; not a memory benchmark | Lightricks repository, 28 Sep 2026 |
| Mochi 1 preview | Reference single-GPU implementation | Approximately 60GB; not a ComfyUI minimum | No fixed frame-count setup | Genmo model card, 28 Sep 2026 |
FLUX VRAM is more than the diffusion checkpoint
Black Forest Labs' dev and schnell model cards, as of 28 September 2026, each describe 12 billion parameters. Multiplying 12 billion by two bytes per BF16 weight gives about 24 GB (derived). At one byte per FP8 weight, that becomes about 12 GB. Neither calculation includes the rest of the graph.
ComfyUI's official FLUX examples, as of 28 September 2026, recommend full 16-bit weights when resources permit because FP8 can degrade quality. They also offer an FP8 T5 encoder. Choose the encoder and diffusion checkpoint together, then compare outputs before accepting the lower-memory path.
Do not budget from a downloaded file's size. For a rental decision, require a completed generation with the exact graph you intend to use. Record peak memory and elapsed time through decoding and saving, not just successful checkpoint loading. Keep that graph as your acceptance test when switching GPU families.
Wan 2.2 VRAM depends on the variant
Wan-Video's official commands and Comfy Org's native tutorial, both as of 28 September 2026, describe different execution paths. The native 5B offloading claim does not replace Wan's CLI requirement. Keep the implementation attached to the number when you choose a template.
For the best GPU for AI video generation, first narrow the question to a model, precision, resolution and frame count. Tencent's original HunyuanVideo requirements, as of 28 September 2026, illustrate why: its reported peak, a VENDOR CLAIM, rises between the two resolutions in the table while frame count stays fixed. Use the larger intended output as your rental acceptance test.
Map memory to the smallest candidate GPU
The site's verified GPU specifications provide the capacity side of the comparison. The table also separates dense and sparse compute figures; neither is a ComfyUI completion-time benchmark.
| Spec | RTX 4090 | RTX 5090 | L40S | RTX PRO 6000 | A100 | H100 |
|---|---|---|---|---|---|---|
| VRAM | 24 GB | 32 GB | 48 GB | 96 GB | 40 to 80 GB | 80 to 94 GB |
| Memory type | GDDR6X | GDDR7 | GDDR6 | GDDR7 | HBM2e | HBM3 |
| Memory bandwidth | 1,008 GB/s | 1,792 GB/s | 864 GB/s | 1,792 GB/s | 2,039 GB/s | 3,350 GB/s |
| FP8 (dense) | 330.3 TFLOPS | 419 TFLOPS | 733 TFLOPS | 1,007.6 TFLOPS | Not published | 1,979 TFLOPS |
| FP8 (with sparsity) | 660.6 TFLOPS | 838 TFLOPS | 1,466 TFLOPS | 2,015.2 TFLOPS | Not published | 3,958 TFLOPS |
| FP4 (dense) | Not published | 1,676 TFLOPS | Not published | 2,015.2 TFLOPS | Not published | Not published |
| FP4 (with sparsity) | Not published | 3,352 TFLOPS | Not published | 4,030.4 TFLOPS | Not published | Not published |
The mapping below is derived capacity screening, restricted to the families compared on this page. It is not a tested fit guarantee or a claim that smaller GPUs elsewhere cannot work. Where the documentation gives no complete runtime requirement, there is no defensible smallest guaranteed card. A row that relies on a figure labelled VENDOR CLAIM or derived in the first table keeps that status here.
| Workflow | Smallest capacity candidate in this comparison | Qualification |
|---|---|---|
| SDXL Base or Base plus Refiner; SD 3.5 Large FP16 | RTX 3090 or RTX 4090, 24GB | Covers stated guidance; validate the full graph |
| SD 3.5 Medium | RTX 3090 or RTX 4090, 24GB, provisional | Encoder memory is excluded from Stability AI's figure |
| FLUX.1 BF16 | RTX 5090, 32GB, first candidate above weight-only estimate | No guaranteed whole-pipeline fit; test or consider L40S for headroom |
| FLUX.1 FP8; HiDream FP8 | RTX 3090 or RTX 4090, 24GB | FLUX remains provisional; HiDream guidance explicitly exceeds 16GB |
| HiDream 16-bit | RTX 5090, 32GB | Above the guide's strict greater-than-27GB guidance |
| Wan 2.1 1.3B; Wan 2.2 native 5B or offloaded CLI 5B | RTX 3090 or RTX 4090, 24GB | Preserve each path's offloading settings |
| Wan 2.1 14B | RTX 3090 or RTX 4090, 24GB, provisional only | "24GB+" does not establish a ceiling |
| HunyuanVideo's 45GB path | L40S, 48GB | Capacity screen only; Tencent recommends an 80GB GPU |
| HunyuanVideo's 60GB path; Mochi reference; Wan 2.2 A14B CLI | A100 80GB variant or H100 80GB variant | A100 40GB is insufficient for these stated requirements |
| LTX-Video | No minimum assigned | Validate the exact checkpoint and implementation |
Start an image shortlist with RTX 4090 rentals and the RTX 3090 versus RTX 4090 comparison. For a GPU for Stable Diffusion, those are candidates to price, not a reason to discard smaller cards that meet SDXL's guidance. The best GPU for Stable Diffusion is the one that completes your accepted graph at the lowest total cost.
Move to RTX 5090 rentals when the capacity screen calls for more memory. Compare L40S rentals and L40S versus RTX 4090 for the next step. The RTX PRO 6000 guide covers the larger option: the site's verified table gives that family 96GB.
Read the live table after filtering for memory, then compare each suitable offer using the billed time your workflow actually needs.
FP8 and FP4 acceleration require the matching path
VENDOR CLAIM: NVIDIA's ComfyUI tutorial, as of 28 September 2026, recommends FP4 models on RTX 50 Series and FP8 on RTX 40 Series. NVIDIA's CES 2026 ComfyUI material announces native NVFP4 and NVFP8 support. Treat this as vendor guidance for supported paths, not a measured speed ranking for your custom graph.
VENDOR CLAIM: NVIDIA's TensorRT support matrix, dated 8 September 2026, explicitly describes FP4 on H100 and L40S as hardware emulation without accelerated FP4 linear operations. That software support label does not establish native acceleration. Choose precision for an acceptable output first, and evaluate the matching workflow before paying for a hardware feature.
Choose the environment before loading models
Comfy Org announced the repository's move into its organization on 29 December 2025. Its repository describes ComfyUI as an open-source node-graph application and inference backend. As of 28 September 2026, GitHub listed Core v0.37.0, released 21 September 2026, as the latest core release. Comfy Org maintains separate Core, Desktop and Frontend release processes; pin the component versions you actually deploy.
Comfy Org's requirements documentation, as of 28 September 2026, lists NVIDIA CUDA, AMD ROCm, Intel Arc through PyTorch torch.xpu, and Apple Silicon through MPS/Metal. Check your chosen environment against the workflow's dependencies before renting.
Comfy Org's CLI help on that date says --lowvram has no effect with dynamic VRAM enabled. With dynamic VRAM disabled, it moves text encoders to the CPU. It describes --novram as the fallback when low-VRAM mode is insufficient, and --cpu as slow CPU-only execution. Do not carry an old launch command into a new template without checking those semantics.
ComfyUI RunPod templates and hosted alternatives
RunPod's template reference, verified 25 August 2026, documents a CUDA 12.8 ComfyUI Pod template and a separate CUDA 13 template for Blackwell. RunPod says they auto-start ComfyUI and begin with an empty checkpoint directory. Budget a setup session to install your intended models and confirm the graph runs before planning a batch.
Vast.ai's 19 December 2025 announcement describes a ComfyUI Serverless template with a JSON-workflow worker and API wrapper. It uploads outputs to S3-compatible storage and returns pre-signed URLs. Vast.ai's Wan 2.2 template documentation, as of 28 September 2026, describes a dedicated T2V A14B template that downloads models during first-boot provisioning. Use the serverless GPU pricing guide when deciding how to cost that deployment pattern.
RunComfy's website, as of 28 September 2026, describes browser access and Cloud Save for workflow JSON, environment, nodes and models. Compare these options on the work you want to own: environment setup, reproducibility, output storage and deployment. Require a successful run of your intended graph before choosing any hosted service.
ComfyUI pricing: subscription versus GPU-hours
Comfy Org describes Comfy Cloud as its official hosted ComfyUI service, accessible without a local GPU. It is a subscription service, not a GPU rental tracked in the live table. Comfy Org published the following dollar-denominated plans on its pricing page as of 28 September 2026.
| Plan | Monthly subscription | Annual billing | Displayed monthly equivalent on annual plan |
|---|---|---|---|
| Standard | $20 | $192 | $16 |
| Creator | $35 | $336 | $28 |
| Pro | $100 | $960 | $80 |
Comfy Org's Comfy Cloud page, as of 28 September 2026, says credits cover active GPU execution and Partner Node API models, while workflow editing is free. It limits Standard and Creator jobs to 30 minutes and Pro jobs to one hour. Its pricing page on that date says subscription credits expire at the billing-cycle reset. Choose a plan against job duration and expected consumption, not just the subscription amount.
For the rental comparison, multiply your expected billed GPU-hours per month by the suitable offer's live rate. Add storage and other applicable charges. Compare that total with the subscription and any required credit top-ups. Do not assume a credit converts into a fixed GPU-hour across every workflow.
Keep active generation time and billed session time separate in your estimate. Include your intended setup and idle time when the rental bills for them. Then divide each route's total monthly cost by the number of accepted images or clips. A cheap failed generation does not help your budget; record retries and rejected outputs during the acceptance run.
The rental decision
Choose the smallest capacity tier supported by your largest workflow's documented requirements. Run that graph once at its intended dimensions and length, then compare total cost using measured billed time. Select Comfy Cloud when its workflow support, job limits and credit budget suit your usage; select a rental when you need control over the environment. Increase capacity when the accepted workflow demands it, and change precision only after checking output quality.
Sources
- Comfy Org repository, accessed 28 September 2026.
- Comfy Org repository move, 29 December 2025.
- ComfyUI Core v0.37.0, 21 September 2026; latest status accessed 28 September 2026.
- Comfy Org system requirements, accessed 28 September 2026.
- Comfy Org CLI argument definitions, accessed 28 September 2026.
- Comfy Org SDXL and SD 3.5 guidance, accessed 28 September 2026.
- Stability AI SD 3.5 announcement, 22 October 2024, updated 29 October 2024 for Medium.
- Black Forest Labs FLUX.1 dev and FLUX.1 schnell, accessed 28 September 2026.
- ComfyUI FLUX examples, accessed 28 September 2026.
- Comfy Org HiDream guide, accessed 28 September 2026.
- Wan 2.1 repository and Comfy Org Wan overview, accessed 28 September 2026.
- Wan 2.2 repository and Comfy Org native Wan 2.2 guide, accessed 28 September 2026.
- Tencent HunyuanVideo requirements, accessed 28 September 2026.
- Lightricks LTX-Video repository, accessed 28 September 2026; resolution update 15 April 2025.
- Genmo Mochi model card, accessed 28 September 2026.
- NVIDIA ComfyUI tutorial and CES 2026 ComfyUI material, accessed 28 September 2026.
- NVIDIA TensorRT support matrix, dated 8 September 2026.
- RunPod ComfyUI template reference, verified 25 August 2026.
- Vast.ai ComfyUI Serverless, 19 December 2025; Vast.ai Wan 2.2 template, accessed 28 September 2026.
- RunComfy, accessed 28 September 2026.
- Comfy Cloud service and limits and Comfy Cloud subscription pricing, accessed 28 September 2026.