ComfyUI Cloud: Rent for Your Largest Model's VRAM

Match SDXL, FLUX and video workflows to GPU memory, compare live rental options, and weigh Comfy Cloud subscriptions against your monthly GPU usage.

By Faiz Ahmed•
•13 min read

For ComfyUI cloud use, image models such as SDXL fit modest cards: Comfy Org's guide, as of 28 September 2026, lists 8GB minimum for SDXL Base FP16. Full 16-bit FLUX.1 needs roughly 24 GB for its diffusion weights alone, derived from the parameter count on Black Forest Labs' model card, while Wan's official A14B video commands specify at least 80GB, so choose by the biggest model and workflow you will load. The cheapest suitable GPUs can then be selected from the live table below using those memory requirements.

ComfyUI GPU requirements start with the workflow

Write down the checkpoint, precision, output dimensions, batch size and video length before opening a rental listing. Treat each figure below as applying only to the named path. A smaller image or an offloaded encoder is a different configuration, not evidence that every graph fits the same card.

Dates in the table identify the publisher's documentation as accessed on that date unless a publication date is stated. "Not specified" means you should not infer a missing condition. The table deliberately distinguishes reported runtime memory from weight-only arithmetic.

ModelVariant or precisionVRAM guidance and limitsResolution or frames statedSource and date
SDXLBase FP168GB minimum; 12GB recommended; batch and offloading unspecifiedNative 1024x1024, not a complete memory test setupComfy Org model guide, 28 Sep 2026
SDXLBase plus Refiner16GB+ guidance; sequential offloading may change fitNot specified for this memory figureComfy Org model guide, 28 Sep 2026
SD 3.5 MediumPrecision unspecifiedVENDOR CLAIM: 9.9 GB excluding text encodersModel supports 0.25 to 2 megapixels; no memory-test resolution givenStability AI announcement, 22 Oct 2024, Medium update 29 Oct 2024
SD 3.5 LargeFP1616 to 24GB; offloading unspecifiedNot specifiedComfy Org model guide, 28 Sep 2026
FLUX.1 dev and schnellBF16 diffusion weights onlyDerived: about 24 GB, excluding encoders, VAE and runtime allocationsdev model-card example: 1024x1024 with CPU offloading; not a peak-memory measurementBlack Forest Labs model cards, 28 Sep 2026
FLUX.1 dev and schnellFP8 diffusion weights onlyDerived: about 12 GB before metadata and runtime allocations; official examples warn of degraded qualityNo complete-workflow minimum establishedBlack Forest Labs and ComfyUI examples, 28 Sep 2026
HiDream-I1Full, Dev or Fast, FP8 / 16-bitMore than 16GB / more than 27GB; FP8 quality equivalence is not established by this fit guidanceNot specifiedComfy Org HiDream guide, 28 Sep 2026
Wan 2.1T2V-1.3BVENDOR CLAIM: 8.19 GB; measurement setup incompleteSeparate official example: 832x480 with model offload and CPU T5Wan-Video repository, 28 Sep 2026
Wan 2.114B, precision unspecified24GB+ broad guidance; not a verified production minimumFrame count and offloading unspecifiedComfy Org Wan overview, 28 Sep 2026
Wan 2.2TI2V-5B official CLIAt least 24GB with model offload, dtype conversion and CPU T51280x704Wan-Video repository, 28 Sep 2026
Wan 2.2TI2V-5B native ComfyUI FP16VENDOR CLAIM: should fit on 8GB with native offloading; not an all-resident fitNo frame-count benchmark attachedComfy Org Wan 2.2 guide, 28 Sep 2026
Wan 2.2T2V-A14B / I2V-A14B official CLIAt least 80GB; T2V command uses offloading and dtype conversion1280x720Wan-Video repository, 28 Sep 2026
Original HunyuanVideoPrecision unspecified in requirements tableVENDOR CLAIM: 45GB / 60GB peak, batch size 1544x960 / 720x1280, both 129 framesTencent Hunyuan repository, 28 Sep 2026
LTX-VideoOriginal 2B / 13B; FP8 checkpoints also publishedNo stock-workflow minimum established here; FP8 quality equivalence unestablishedLightricks' 15 Apr 2025 default: 1216x704 at 30 FPS; not a memory benchmarkLightricks repository, 28 Sep 2026
Mochi 1 previewReference single-GPU implementationApproximately 60GB; not a ComfyUI minimumNo fixed frame-count setupGenmo model card, 28 Sep 2026

FLUX VRAM is more than the diffusion checkpoint

Black Forest Labs' dev and schnell model cards, as of 28 September 2026, each describe 12 billion parameters. Multiplying 12 billion by two bytes per BF16 weight gives about 24 GB (derived). At one byte per FP8 weight, that becomes about 12 GB. Neither calculation includes the rest of the graph.

ComfyUI's official FLUX examples, as of 28 September 2026, recommend full 16-bit weights when resources permit because FP8 can degrade quality. They also offer an FP8 T5 encoder. Choose the encoder and diffusion checkpoint together, then compare outputs before accepting the lower-memory path.

Do not budget from a downloaded file's size. For a rental decision, require a completed generation with the exact graph you intend to use. Record peak memory and elapsed time through decoding and saving, not just successful checkpoint loading. Keep that graph as your acceptance test when switching GPU families.

Wan 2.2 VRAM depends on the variant

Wan-Video's official commands and Comfy Org's native tutorial, both as of 28 September 2026, describe different execution paths. The native 5B offloading claim does not replace Wan's CLI requirement. Keep the implementation attached to the number when you choose a template.

For the best GPU for AI video generation, first narrow the question to a model, precision, resolution and frame count. Tencent's original HunyuanVideo requirements, as of 28 September 2026, illustrate why: its reported peak, a VENDOR CLAIM, rises between the two resolutions in the table while frame count stays fixed. Use the larger intended output as your rental acceptance test.

Map memory to the smallest candidate GPU

The site's verified GPU specifications provide the capacity side of the comparison. The table also separates dense and sparse compute figures; neither is a ComfyUI completion-time benchmark.

SpecRTX 4090RTX 5090L40SRTX PRO 6000A100H100
VRAM24 GB32 GB48 GB96 GB40 to 80 GB80 to 94 GB
Memory typeGDDR6XGDDR7GDDR6GDDR7HBM2eHBM3
Memory bandwidth1,008 GB/s1,792 GB/s864 GB/s1,792 GB/s2,039 GB/s3,350 GB/s
FP8 (dense)330.3 TFLOPS419 TFLOPS733 TFLOPS1,007.6 TFLOPSNot published1,979 TFLOPS
FP8 (with sparsity)660.6 TFLOPS838 TFLOPS1,466 TFLOPS2,015.2 TFLOPSNot published3,958 TFLOPS
FP4 (dense)Not published1,676 TFLOPSNot published2,015.2 TFLOPSNot publishedNot published
FP4 (with sparsity)Not published3,352 TFLOPSNot published4,030.4 TFLOPSNot publishedNot published
Figures from the vendor datasheets: RTX 4090, RTX 5090, L40S, RTX PRO 6000, A100, H100, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.

The mapping below is derived capacity screening, restricted to the families compared on this page. It is not a tested fit guarantee or a claim that smaller GPUs elsewhere cannot work. Where the documentation gives no complete runtime requirement, there is no defensible smallest guaranteed card. A row that relies on a figure labelled VENDOR CLAIM or derived in the first table keeps that status here.

WorkflowSmallest capacity candidate in this comparisonQualification
SDXL Base or Base plus Refiner; SD 3.5 Large FP16RTX 3090 or RTX 4090, 24GBCovers stated guidance; validate the full graph
SD 3.5 MediumRTX 3090 or RTX 4090, 24GB, provisionalEncoder memory is excluded from Stability AI's figure
FLUX.1 BF16RTX 5090, 32GB, first candidate above weight-only estimateNo guaranteed whole-pipeline fit; test or consider L40S for headroom
FLUX.1 FP8; HiDream FP8RTX 3090 or RTX 4090, 24GBFLUX remains provisional; HiDream guidance explicitly exceeds 16GB
HiDream 16-bitRTX 5090, 32GBAbove the guide's strict greater-than-27GB guidance
Wan 2.1 1.3B; Wan 2.2 native 5B or offloaded CLI 5BRTX 3090 or RTX 4090, 24GBPreserve each path's offloading settings
Wan 2.1 14BRTX 3090 or RTX 4090, 24GB, provisional only"24GB+" does not establish a ceiling
HunyuanVideo's 45GB pathL40S, 48GBCapacity screen only; Tencent recommends an 80GB GPU
HunyuanVideo's 60GB path; Mochi reference; Wan 2.2 A14B CLIA100 80GB variant or H100 80GB variantA100 40GB is insufficient for these stated requirements
LTX-VideoNo minimum assignedValidate the exact checkpoint and implementation

Start an image shortlist with RTX 4090 rentals and the RTX 3090 versus RTX 4090 comparison. For a GPU for Stable Diffusion, those are candidates to price, not a reason to discard smaller cards that meet SDXL's guidance. The best GPU for Stable Diffusion is the one that completes your accepted graph at the lowest total cost.

Move to RTX 5090 rentals when the capacity screen calls for more memory. Compare L40S rentals and L40S versus RTX 4090 for the next step. The RTX PRO 6000 guide covers the larger option: the site's verified table gives that family 96GB.

GPUCheapest $/GPU-hrProviderProviders in stock
RTX 3090$0.27Vast.ai3
RTX 4090$0.53Vast.ai3
RTX 5090$0.76Vast.ai3
L40S$0.80Vast.ai4
RTX PRO 6000 Blackwell$0.59RunPod6
H100$2.59Vast.ai7
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: .

Read the live table after filtering for memory, then compare each suitable offer using the billed time your workflow actually needs.

FP8 and FP4 acceleration require the matching path

VENDOR CLAIM: NVIDIA's ComfyUI tutorial, as of 28 September 2026, recommends FP4 models on RTX 50 Series and FP8 on RTX 40 Series. NVIDIA's CES 2026 ComfyUI material announces native NVFP4 and NVFP8 support. Treat this as vendor guidance for supported paths, not a measured speed ranking for your custom graph.

VENDOR CLAIM: NVIDIA's TensorRT support matrix, dated 8 September 2026, explicitly describes FP4 on H100 and L40S as hardware emulation without accelerated FP4 linear operations. That software support label does not establish native acceleration. Choose precision for an acceptable output first, and evaluate the matching workflow before paying for a hardware feature.

Choose the environment before loading models

Comfy Org announced the repository's move into its organization on 29 December 2025. Its repository describes ComfyUI as an open-source node-graph application and inference backend. As of 28 September 2026, GitHub listed Core v0.37.0, released 21 September 2026, as the latest core release. Comfy Org maintains separate Core, Desktop and Frontend release processes; pin the component versions you actually deploy.

Comfy Org's requirements documentation, as of 28 September 2026, lists NVIDIA CUDA, AMD ROCm, Intel Arc through PyTorch torch.xpu, and Apple Silicon through MPS/Metal. Check your chosen environment against the workflow's dependencies before renting.

Comfy Org's CLI help on that date says --lowvram has no effect with dynamic VRAM enabled. With dynamic VRAM disabled, it moves text encoders to the CPU. It describes --novram as the fallback when low-VRAM mode is insufficient, and --cpu as slow CPU-only execution. Do not carry an old launch command into a new template without checking those semantics.

ComfyUI RunPod templates and hosted alternatives

RunPod's template reference, verified 25 August 2026, documents a CUDA 12.8 ComfyUI Pod template and a separate CUDA 13 template for Blackwell. RunPod says they auto-start ComfyUI and begin with an empty checkpoint directory. Budget a setup session to install your intended models and confirm the graph runs before planning a batch.

Vast.ai's 19 December 2025 announcement describes a ComfyUI Serverless template with a JSON-workflow worker and API wrapper. It uploads outputs to S3-compatible storage and returns pre-signed URLs. Vast.ai's Wan 2.2 template documentation, as of 28 September 2026, describes a dedicated T2V A14B template that downloads models during first-boot provisioning. Use the serverless GPU pricing guide when deciding how to cost that deployment pattern.

RunComfy's website, as of 28 September 2026, describes browser access and Cloud Save for workflow JSON, environment, nodes and models. Compare these options on the work you want to own: environment setup, reproducibility, output storage and deployment. Require a successful run of your intended graph before choosing any hosted service.

ComfyUI pricing: subscription versus GPU-hours

Comfy Org describes Comfy Cloud as its official hosted ComfyUI service, accessible without a local GPU. It is a subscription service, not a GPU rental tracked in the live table. Comfy Org published the following dollar-denominated plans on its pricing page as of 28 September 2026.

PlanMonthly subscriptionAnnual billingDisplayed monthly equivalent on annual plan
Standard$20$192$16
Creator$35$336$28
Pro$100$960$80

Comfy Org's Comfy Cloud page, as of 28 September 2026, says credits cover active GPU execution and Partner Node API models, while workflow editing is free. It limits Standard and Creator jobs to 30 minutes and Pro jobs to one hour. Its pricing page on that date says subscription credits expire at the billing-cycle reset. Choose a plan against job duration and expected consumption, not just the subscription amount.

For the rental comparison, multiply your expected billed GPU-hours per month by the suitable offer's live rate. Add storage and other applicable charges. Compare that total with the subscription and any required credit top-ups. Do not assume a credit converts into a fixed GPU-hour across every workflow.

Keep active generation time and billed session time separate in your estimate. Include your intended setup and idle time when the rental bills for them. Then divide each route's total monthly cost by the number of accepted images or clips. A cheap failed generation does not help your budget; record retries and rejected outputs during the acceptance run.

The rental decision

Choose the smallest capacity tier supported by your largest workflow's documented requirements. Run that graph once at its intended dimensions and length, then compare total cost using measured billed time. Select Comfy Cloud when its workflow support, job limits and credit budget suit your usage; select a rental when you need control over the environment. Increase capacity when the accepted workflow demands it, and change precision only after checking output quality.

Sources

Frequently asked questions

Which GPU should I rent for Stable Diffusion?▾

For SDXL Base, Comfy Org's guide gives 8GB minimum and 12GB recommended for FP16. Start with the smallest suitable rental and validate your complete workflow before committing.

How much VRAM does FLUX.1 need?▾

Its 12 billion parameters at two bytes each put the BF16 diffusion weights alone at about 24 GB (derived), before encoders, VAE and runtime memory. ComfyUI's official examples offer FP8 checkpoints but warn that output quality can degrade.

Does Wan 2.2 fit on an RTX 4090?▾

Wan's official TI2V-5B command specifies at least 24GB with offloading and T5 on the CPU. Its A14B command instead specifies at least 80GB; the model variant matters.

Is Comfy Cloud the same as renting a GPU?▾

Comfy Cloud is Comfy Org's hosted service with subscription credits. Compare its workflow limits and effective cost per accepted output with a rental's billed GPU-hours.

Should I enable ComfyUI lowvram?▾

Comfy Org's CLI help, as of 28 September 2026, says --lowvram has no effect while dynamic VRAM is enabled. With dynamic VRAM disabled, it moves text encoders to the CPU.

Related Posts