LoRA vs QLoRA: Choose by VRAM, Then Check Quality

Compare LoRA, QLoRA and full fine-tuning with worked 7B and 70B memory budgets, published quality results, GPU choices and live prices to estimate run costs.

By Faiz Ahmed•
•11 min read

Machine-learning LoRA trains low-rank adapters over frozen weights; QLoRA adds a frozen 4-bit base, following Hu and colleagues' 2021 LoRA paper and Dettmers and colleagues' 2023 QLoRA paper. LoRA removes gradient and optimizer storage for the frozen base model, while QLoRA also shrinks base-weight storage. Choose QLoRA when memory is the constraint, LoRA when you have the memory, and full fine-tuning only when adapters fall short and you can pay for the cluster.

This is LoRA for machine learning, not LoRa radio. The rental decision starts with what you train, then what remains in memory during a step. Use the training versus inference guide if you are budgeting both adaptation and serving.

LoRA vs QLoRA changes the memory bill

The comparison combines Hu et al. (October 2021), Dettmers et al. (May 2023), Biderman et al. (September 2024), and Hugging Face and Axolotl documentation accessed 28 September 2026. Precision entries describe the configurations budgeted below.

MethodWhat is trainedBase-model precisionMemory componentsQuality evidenceDocumented tooling
Full fine-tuningAll model weightsMixed precision, with a master copyWeights, all gradients and optimizer states, activations, buffersBiderman et al.: stronger continued-pretraining results in their studyAxolotl
LoRALow-rank adapters; base frozen16-bit baseBase weights, adapter weights/gradients/states, activations, buffersBiderman et al.: higher ranks can match full instruction tuningPEFT, Axolotl
QLoRALow-rank adapters; quantized base frozen4-bit storage, higher-precision computeQuantized base and scales, adapter states, activations, buffersDettmers et al.: near-identical aggregate MMLU to BF16 LoRAPEFT, Axolotl

Hu's LoRA preprint, submitted 17 June 2021, lists Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen. Its October 2021 revision trains low-rank update matrices A and B on top of frozen pretrained weights. Derived: for a d-by-k weight and rank r, the adapter count is r(d+k), rather than d×k trainable elements.

VENDOR CLAIM: Microsoft's LoRA paper, in its October 2021 revision, reports 10,000 times fewer trainable parameters and a threefold reduction in GPU-memory requirements against GPT-3 175B full fine-tuning with Adam. Those reductions describe that experiment, not every LoRA fine tuning job.

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer submitted QLoRA on 23 May 2023. Their headline claim includes these exact words: "finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance".

Their paper describes 4-bit NormalFloat (NF4) for normally distributed weights, double quantization of quantization constants, and paged optimizers to handle memory spikes through NVIDIA unified memory. It reports scaling metadata falling from 0.5 to 0.127 bits per parameter. Its experiments use BF16 computation, rank 64, alpha 16 and adapters on all linear layers. Read its results with those settings in mind, not just the method's name.

How much VRAM to fine tune: six worked budgets

Derived: use exactly 7 billion or 70 billion base parameters, called P. For the adapter examples, choose T = 1% of P as an illustrative planning assumption, not a published architecture or recommended rank. Count decimal GB throughout. These calculations estimate model states before activations, temporary allocations and runtime overhead.

Hugging Face's memory anatomy documentation, accessed 28 September 2026, gives mixed-precision weights at 6 bytes per trainable parameter, FP32 gradients at 4 bytes and Adam states at 8 bytes. That makes 18P for full tuning. Apply the same conservative mixed-precision accounting to adapters: 18T, comprising adapter weights and master copy, gradients and Adam states.

For the frozen LoRA base, use 2P bytes, consistent with NVIDIA's 1 September 2026 FP16/BF16 storage guidance. For QLoRA, assume every base parameter receives the QLoRA paper's 4-bit storage plus 0.127-bit scaling overhead: 4.127P/8 bytes. This idealized coverage is an explicit assumption. The precision guide explains the formats separately.

MethodDerived 7B example, T = 70 millionDerived 70B example, T = 700 million
Full mixed-precision Adam7 × 18 = 126 GB70 × 18 = 1,260 GB
LoRA, 16-bit frozen base7 × 2 + 0.07 × 18 = 15.26 GB70 × 2 + 0.7 × 18 = 152.6 GB
QLoRA, ideal NF4 + double quantization7 × 4.127/8 + 0.07 × 18 = about 4.87 GB70 × 4.127/8 + 0.7 × 18 = about 48.71 GB

Sequence length, batch size and gradient checkpointing are unspecified in this state-only accounting. They are not secretly set to favorable values. Do not read these numbers as peak VRAM. Hugging Face's documentation notes that implementations materializing attention-score matrices have storage growing quadratically with sequence length.

The ZeRO paper, dated 13 May 2020, instead assumes FP16 gradients and reaches 16 bytes per parameter. Derived under its assumptions, full-tuning states become 112 GB for 7B and 1,120 GB for 70B. The difference is precision accounting, not a contradiction. ZeRO also excludes activations, buffers and fragmentation from that state total.

Torchtune's QLoRA tutorial, accessed 28 September 2026, keeps adapters, activations, gradients and optimizer states at higher precision. Its documented builder quantizes only layers configured with adapters. Consequently, actual unquantized layers can push your budget above the ideal QLoRA arithmetic. Enter your configuration in the LLM VRAM calculator, then verify it with a short training run.

Published peaks and the current recipe settings

VENDOR CLAIM: torchtune's Llama2-7B tutorial, accessed 28 September 2026, reports these peak reserved allocations. The tutorial does not identify a GPU model or immutable benchmark configuration.

Published comparisonLoRAQLoRA
Torchtune Llama2-7B peak reserved memory15.57 GB9.29 GB

Torchtune's linked recipe, as accessed on 28 September 2026, specifies one CUDA device, BF16 compute, batch size 2, eight accumulation steps, activation checkpointing enabled and activation offloading disabled. It uses rank 8, alpha 16, fused AdamW, and query, value, attention-output and MLP adapters. The dataset is unpacked alpaca_cleaned_dataset examples; max_seq_len=null supplies no fixed numeric sequence cap. These mutable recipe settings are not verified historical settings for those peaks.

VENDOR CLAIM: Hugging Face PEFT's table, accessed 28 September 2026, uses twitter_complaints, an A100 80 GB and more than 64 GB of host RAM. Its T0_3B entries are reproduced below.

MethodGPU memoryCPU memory
Full fine-tuning47.14 GB2.96 GB
LoRA14.4 GB2.96 GB
LoRA with DeepSpeed CPU offloading9.8 GB17.8 GB

PEFT does not state sequence length, batch size, precision or gradient-checkpointing settings for that table. It demonstrates the CPU-memory trade, but cannot settle whether your own recipe fits. Use the GPU requirements reference for another planning check.

Quality decides whether adapters are enough

Dettmers et al.'s May 2023 Table 4 reports mean five-shot MMLU accuracy of 53.1 for NF4 plus double quantization versus 53.0 for BF16 LoRA across LLaMA 7B to 65B Alpaca/FLAN-v2 runs. The large-model baseline is LoRA. Direct comparisons against full fine-tuning use smaller RoBERTa/T5 models. Do not turn that result into universal full-tuning equivalence.

Dan Biderman and colleagues' LoRA Learns Less and Forgets Less, revised 20 September 2024, studied Llama-2-7B adaptation for programming and mathematics. They report weaker target-domain learning with standard low ranks, but better retention outside the target domain. Higher ranks could match full fine-tuning for instruction tuning; they did not close the continued-pretraining gap. Their main experiments used LionW, not AdamW.

Set a target-task score and a retention check before renting. Compare LoRA and QLoRA on the same evaluation set. If both fail, test broader adapter coverage or higher rank before committing to full tuning. Recalculate T whenever you change coverage or rank.

Choose tooling before choosing the card

Hugging Face's PEFT documentation, accessed 28 September 2026, provides adapter configuration; its quantization guide uses bitsandbytes NF4, double quantization and BF16 compute, with prepare_model_for_kbit_training() before adapter configuration. Its LoRA API supports target_modules="all-linear", excluding the output layer for a PreTrainedModel. Hugging Face's TRL documentation on the same date provides SFTTrainer for supervised fine-tuning.

Axolotl's documentation, accessed 28 September 2026, supports all three methods through YAML configurations and documents FSDP2 and DeepSpeed for multiple GPUs. Torchtune's documentation on that date provides native-PyTorch recipes with checkpointing and accumulation. LLaMA-Factory's README on that date offers CLI and web interfaces. Choose the interface your team can reproduce and inspect.

VENDOR CLAIM: Unsloth's guide, accessed 28 September 2026, supports LoRA, QLoRA and full fine-tuning, and recommends QLoRA as an accessible starting method. That guidance does not establish a fixed speedup or job-cost saving.

Fine tuning GPU requirements and rental candidates

GPUperhour's datasheet-verified NVIDIA family specifications supply the capacities below. Read capacity first; use the HBM versus GDDR guide when comparing the memory columns.

SpecRTX 4090RTX 5090L40SRTX PRO 6000A100H100H200
VRAM24 GB32 GB48 GB96 GB40 to 80 GB80 to 94 GB141 GB
Memory typeGDDR6XGDDR7GDDR6GDDR7HBM2eHBM3HBM3e
Memory bandwidth1,008 GB/s1,792 GB/s864 GB/s1,792 GB/s2,039 GB/s3,350 GB/s4,800 GB/s
FP16 (dense)165.2 TFLOPS209.5 TFLOPS362.05 TFLOPS503.8 TFLOPS312 TFLOPS989 TFLOPS989 TFLOPS
FP16 (with sparsity)330.4 TFLOPS419 TFLOPS733 TFLOPS1,007.6 TFLOPS624 TFLOPS1,979 TFLOPS1,979 TFLOPS
FP8 (dense)330.3 TFLOPS419 TFLOPS733 TFLOPS1,007.6 TFLOPSNot published1,979 TFLOPS1,979 TFLOPS
FP8 (with sparsity)660.6 TFLOPS838 TFLOPS1,466 TFLOPS2,015.2 TFLOPSNot published3,958 TFLOPS3,958 TFLOPS
Figures from the vendor datasheets: RTX 4090, RTX 5090, L40S, RTX PRO 6000, A100, H100, H200, checked 13 Sep 2026. With-sparsity figures assume 2:4 structured sparsity and are twice the dense figure, so compare dense with dense. "Not published" means the vendor gives no figure.
GPUCheapest $/GPU-hrProviderProviders in stock
RTX 4090$0.53Vast.ai3
RTX 5090$0.53Vast.ai2
L40S$0.97Massed Compute4
RTX PRO 6000 Blackwell$0.59RunPod3
A100$0.67Vast.ai9
H100$2.50Hyperstack7
H200$3.43QuantaCloud6
Cheapest in-stock on-demand price per GPU-hour, from providers with live stock tracking. Latest stock observation: . QuantaCloud operates this site and is ranked by price like every other provider.

Derived: these are the smallest capacities in this comparison that clear the modeled states. They are candidates, not guaranteed training fits. The actual smallest card that fits cannot be established without the missing activation budget.

Model and methodState budgetSmallest candidate in this comparison
7B QLoRA4.87 GBRTX 4090, 24 GB
7B LoRA15.26 GBRTX 4090, 24 GB
7B full tuning126 GBH200, 141 GB, only if remaining allocations fit
70B QLoRA48.71 GBA100 80 GB or H100 80 GB
70B LoRA152.6 GBNo single listed card; use sharding or offloading
70B full tuning1,260 GBCluster planning required

The RTX 5090's 32 GB and L40S's 48 GB provide intermediate options if the 7B peak exceeds the RTX 4090. The modeled 70B QLoRA states already exceed the L40S. RTX PRO 6000 Blackwell's 96 GB provides another capacity option above the 80 GB candidates. Compare measured runtime before paying for more capacity than your run uses.

For a cluster, ZeRO's May 2020 paper gives ideal fully partitioned state storage as total states divided by device count. That is not a promise that every allocation divides evenly. Plan the parallelism strategy before booking multiple GPUs.

Turn a short trial into a run budget

Estimate compute cost as GPU count × elapsed hours × live per-GPU-hour price. An illustrative single-GPU run lasting six hours consumes six GPU-hours; four GPUs for six hours consume 24 GPU-hours. Multiply those derived totals by the matching live table entry. No completion time for your dataset is established here.

Time representative training steps, then include evaluation, checkpoint saves and setup time in the rental window. Torchtune's memory-optimization documentation, accessed 28 September 2026, explains that checkpointing recomputes activations and offloading moves them to CPU. Lower peak VRAM therefore does not mean lower total cost; compare measured completion times for the same quality target.

Start with QLoRA if memory blocks the model you need. Use LoRA when its measured peak fits your budget. Pay for full fine-tuning only after adapters miss your quality target and the improvement justifies the cluster bill.

Sources

Frequently asked questions

What is LoRA in machine learning?▾

LoRA freezes pretrained model weights and trains low-rank adapters, as described by Hu and colleagues in 2021. This article concerns machine-learning LoRA, not LoRa radio.

How do LoRA and QLoRA differ?▾

QLoRA adds a frozen 4-bit base to adapter training, reducing base-weight storage. Dettmers and colleagues' 2023 paper uses higher-precision computation, so training memory is not entirely 4-bit.

How much VRAM does a 7B model need for fine-tuning?▾

The derived examples give 126 GB for full mixed-precision Adam model states, 15.26 GB for LoRA and about 4.87 GB for QLoRA. These exclude activations and runtime overhead, and the adapter examples assume trainable parameters equal to 1% of the base count.

Can I fine-tune a 70B model on an L40S?▾

Do not assume it fits: the illustrative QLoRA state budget is already about 48.71 GB, above the L40S's 48 GB. A smaller adapter configuration could change that budget, but needs a measured peak-memory check.

Is QLoRA always cheaper than LoRA?▾

Not necessarily: lower memory does not mean lower job cost. Multiply measured GPU-hours for your recipe by the live rental rate and compare runs that meet the same quality target.

Related Posts