Machine-learning LoRA trains low-rank adapters over frozen weights; QLoRA adds a frozen 4-bit base, following Hu and colleagues' 2021 LoRA paper and Dettmers and colleagues' 2023 QLoRA paper. LoRA removes gradient and optimizer storage for the frozen base model, while QLoRA also shrinks base-weight storage. Choose QLoRA when memory is the constraint, LoRA when you have the memory, and full fine-tuning only when adapters fall short and you can pay for the cluster.
This is LoRA for machine learning, not LoRa radio. The rental decision starts with what you train, then what remains in memory during a step. Use the training versus inference guide if you are budgeting both adaptation and serving.
LoRA vs QLoRA changes the memory bill
The comparison combines Hu et al. (October 2021), Dettmers et al. (May 2023), Biderman et al. (September 2024), and Hugging Face and Axolotl documentation accessed 28 September 2026. Precision entries describe the configurations budgeted below.
| Method | What is trained | Base-model precision | Memory components | Quality evidence | Documented tooling |
|---|---|---|---|---|---|
| Full fine-tuning | All model weights | Mixed precision, with a master copy | Weights, all gradients and optimizer states, activations, buffers | Biderman et al.: stronger continued-pretraining results in their study | Axolotl |
| LoRA | Low-rank adapters; base frozen | 16-bit base | Base weights, adapter weights/gradients/states, activations, buffers | Biderman et al.: higher ranks can match full instruction tuning | PEFT, Axolotl |
| QLoRA | Low-rank adapters; quantized base frozen | 4-bit storage, higher-precision compute | Quantized base and scales, adapter states, activations, buffers | Dettmers et al.: near-identical aggregate MMLU to BF16 LoRA | PEFT, Axolotl |
Hu's LoRA preprint, submitted 17 June 2021, lists Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen. Its October 2021 revision trains low-rank update matrices A and B on top of frozen pretrained weights. Derived: for a d-by-k weight and rank r, the adapter count is r(d+k), rather than d×k trainable elements.
VENDOR CLAIM: Microsoft's LoRA paper, in its October 2021 revision, reports 10,000 times fewer trainable parameters and a threefold reduction in GPU-memory requirements against GPT-3 175B full fine-tuning with Adam. Those reductions describe that experiment, not every LoRA fine tuning job.
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer submitted QLoRA on 23 May 2023. Their headline claim includes these exact words: "finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance".
Their paper describes 4-bit NormalFloat (NF4) for normally distributed weights, double quantization of quantization constants, and paged optimizers to handle memory spikes through NVIDIA unified memory. It reports scaling metadata falling from 0.5 to 0.127 bits per parameter. Its experiments use BF16 computation, rank 64, alpha 16 and adapters on all linear layers. Read its results with those settings in mind, not just the method's name.
How much VRAM to fine tune: six worked budgets
Derived: use exactly 7 billion or 70 billion base parameters, called P. For the adapter examples, choose T = 1% of P as an illustrative planning assumption, not a published architecture or recommended rank. Count decimal GB throughout. These calculations estimate model states before activations, temporary allocations and runtime overhead.
Hugging Face's memory anatomy documentation, accessed 28 September 2026, gives mixed-precision weights at 6 bytes per trainable parameter, FP32 gradients at 4 bytes and Adam states at 8 bytes. That makes 18P for full tuning. Apply the same conservative mixed-precision accounting to adapters: 18T, comprising adapter weights and master copy, gradients and Adam states.
For the frozen LoRA base, use 2P bytes, consistent with NVIDIA's 1 September 2026 FP16/BF16 storage guidance. For QLoRA, assume every base parameter receives the QLoRA paper's 4-bit storage plus 0.127-bit scaling overhead: 4.127P/8 bytes. This idealized coverage is an explicit assumption. The precision guide explains the formats separately.
| Method | Derived 7B example, T = 70 million | Derived 70B example, T = 700 million |
|---|---|---|
| Full mixed-precision Adam | 7 × 18 = 126 GB | 70 × 18 = 1,260 GB |
| LoRA, 16-bit frozen base | 7 × 2 + 0.07 × 18 = 15.26 GB | 70 × 2 + 0.7 × 18 = 152.6 GB |
| QLoRA, ideal NF4 + double quantization | 7 × 4.127/8 + 0.07 × 18 = about 4.87 GB | 70 × 4.127/8 + 0.7 × 18 = about 48.71 GB |
Sequence length, batch size and gradient checkpointing are unspecified in this state-only accounting. They are not secretly set to favorable values. Do not read these numbers as peak VRAM. Hugging Face's documentation notes that implementations materializing attention-score matrices have storage growing quadratically with sequence length.
The ZeRO paper, dated 13 May 2020, instead assumes FP16 gradients and reaches 16 bytes per parameter. Derived under its assumptions, full-tuning states become 112 GB for 7B and 1,120 GB for 70B. The difference is precision accounting, not a contradiction. ZeRO also excludes activations, buffers and fragmentation from that state total.
Torchtune's QLoRA tutorial, accessed 28 September 2026, keeps adapters, activations, gradients and optimizer states at higher precision. Its documented builder quantizes only layers configured with adapters. Consequently, actual unquantized layers can push your budget above the ideal QLoRA arithmetic. Enter your configuration in the LLM VRAM calculator, then verify it with a short training run.
Published peaks and the current recipe settings
VENDOR CLAIM: torchtune's Llama2-7B tutorial, accessed 28 September 2026, reports these peak reserved allocations. The tutorial does not identify a GPU model or immutable benchmark configuration.
| Published comparison | LoRA | QLoRA |
|---|---|---|
| Torchtune Llama2-7B peak reserved memory | 15.57 GB | 9.29 GB |
Torchtune's linked recipe, as accessed on 28 September 2026, specifies one CUDA device, BF16 compute, batch size 2, eight accumulation steps, activation checkpointing enabled and activation offloading disabled. It uses rank 8, alpha 16, fused AdamW, and query, value, attention-output and MLP adapters. The dataset is unpacked alpaca_cleaned_dataset examples; max_seq_len=null supplies no fixed numeric sequence cap. These mutable recipe settings are not verified historical settings for those peaks.
VENDOR CLAIM: Hugging Face PEFT's table, accessed 28 September 2026, uses twitter_complaints, an A100 80 GB and more than 64 GB of host RAM. Its T0_3B entries are reproduced below.
| Method | GPU memory | CPU memory |
|---|---|---|
| Full fine-tuning | 47.14 GB | 2.96 GB |
| LoRA | 14.4 GB | 2.96 GB |
| LoRA with DeepSpeed CPU offloading | 9.8 GB | 17.8 GB |
PEFT does not state sequence length, batch size, precision or gradient-checkpointing settings for that table. It demonstrates the CPU-memory trade, but cannot settle whether your own recipe fits. Use the GPU requirements reference for another planning check.
Quality decides whether adapters are enough
Dettmers et al.'s May 2023 Table 4 reports mean five-shot MMLU accuracy of 53.1 for NF4 plus double quantization versus 53.0 for BF16 LoRA across LLaMA 7B to 65B Alpaca/FLAN-v2 runs. The large-model baseline is LoRA. Direct comparisons against full fine-tuning use smaller RoBERTa/T5 models. Do not turn that result into universal full-tuning equivalence.
Dan Biderman and colleagues' LoRA Learns Less and Forgets Less, revised 20 September 2024, studied Llama-2-7B adaptation for programming and mathematics. They report weaker target-domain learning with standard low ranks, but better retention outside the target domain. Higher ranks could match full fine-tuning for instruction tuning; they did not close the continued-pretraining gap. Their main experiments used LionW, not AdamW.
Set a target-task score and a retention check before renting. Compare LoRA and QLoRA on the same evaluation set. If both fail, test broader adapter coverage or higher rank before committing to full tuning. Recalculate T whenever you change coverage or rank.
Choose tooling before choosing the card
Hugging Face's PEFT documentation, accessed 28 September 2026, provides adapter configuration; its quantization guide uses bitsandbytes NF4, double quantization and BF16 compute, with prepare_model_for_kbit_training() before adapter configuration. Its LoRA API supports target_modules="all-linear", excluding the output layer for a PreTrainedModel. Hugging Face's TRL documentation on the same date provides SFTTrainer for supervised fine-tuning.
Axolotl's documentation, accessed 28 September 2026, supports all three methods through YAML configurations and documents FSDP2 and DeepSpeed for multiple GPUs. Torchtune's documentation on that date provides native-PyTorch recipes with checkpointing and accumulation. LLaMA-Factory's README on that date offers CLI and web interfaces. Choose the interface your team can reproduce and inspect.
VENDOR CLAIM: Unsloth's guide, accessed 28 September 2026, supports LoRA, QLoRA and full fine-tuning, and recommends QLoRA as an accessible starting method. That guidance does not establish a fixed speedup or job-cost saving.
Fine tuning GPU requirements and rental candidates
GPUperhour's datasheet-verified NVIDIA family specifications supply the capacities below. Read capacity first; use the HBM versus GDDR guide when comparing the memory columns.
| Spec | RTX 4090 | RTX 5090 | L40S | RTX PRO 6000 | A100 | H100 | H200 |
|---|---|---|---|---|---|---|---|
| VRAM | 24 GB | 32 GB | 48 GB | 96 GB | 40 to 80 GB | 80 to 94 GB | 141 GB |
| Memory type | GDDR6X | GDDR7 | GDDR6 | GDDR7 | HBM2e | HBM3 | HBM3e |
| Memory bandwidth | 1,008 GB/s | 1,792 GB/s | 864 GB/s | 1,792 GB/s | 2,039 GB/s | 3,350 GB/s | 4,800 GB/s |
| FP16 (dense) | 165.2 TFLOPS | 209.5 TFLOPS | 362.05 TFLOPS | 503.8 TFLOPS | 312 TFLOPS | 989 TFLOPS | 989 TFLOPS |
| FP16 (with sparsity) | 330.4 TFLOPS | 419 TFLOPS | 733 TFLOPS | 1,007.6 TFLOPS | 624 TFLOPS | 1,979 TFLOPS | 1,979 TFLOPS |
| FP8 (dense) | 330.3 TFLOPS | 419 TFLOPS | 733 TFLOPS | 1,007.6 TFLOPS | Not published | 1,979 TFLOPS | 1,979 TFLOPS |
| FP8 (with sparsity) | 660.6 TFLOPS | 838 TFLOPS | 1,466 TFLOPS | 2,015.2 TFLOPS | Not published | 3,958 TFLOPS | 3,958 TFLOPS |
| GPU | Cheapest $/GPU-hr | Provider | Providers in stock |
|---|---|---|---|
| RTX 4090 | $0.53 | Vast.ai | 3 |
| RTX 5090 | $0.53 | Vast.ai | 2 |
| L40S | $0.97 | Massed Compute | 4 |
| RTX PRO 6000 Blackwell | $0.59 | RunPod | 3 |
| A100 | $0.67 | Vast.ai | 9 |
| H100 | $2.50 | Hyperstack | 7 |
| H200 | $3.43 | QuantaCloud | 6 |
Derived: these are the smallest capacities in this comparison that clear the modeled states. They are candidates, not guaranteed training fits. The actual smallest card that fits cannot be established without the missing activation budget.
| Model and method | State budget | Smallest candidate in this comparison |
|---|---|---|
| 7B QLoRA | 4.87 GB | RTX 4090, 24 GB |
| 7B LoRA | 15.26 GB | RTX 4090, 24 GB |
| 7B full tuning | 126 GB | H200, 141 GB, only if remaining allocations fit |
| 70B QLoRA | 48.71 GB | A100 80 GB or H100 80 GB |
| 70B LoRA | 152.6 GB | No single listed card; use sharding or offloading |
| 70B full tuning | 1,260 GB | Cluster planning required |
The RTX 5090's 32 GB and L40S's 48 GB provide intermediate options if the 7B peak exceeds the RTX 4090. The modeled 70B QLoRA states already exceed the L40S. RTX PRO 6000 Blackwell's 96 GB provides another capacity option above the 80 GB candidates. Compare measured runtime before paying for more capacity than your run uses.
For a cluster, ZeRO's May 2020 paper gives ideal fully partitioned state storage as total states divided by device count. That is not a promise that every allocation divides evenly. Plan the parallelism strategy before booking multiple GPUs.
Turn a short trial into a run budget
Estimate compute cost as GPU count × elapsed hours × live per-GPU-hour price. An illustrative single-GPU run lasting six hours consumes six GPU-hours; four GPUs for six hours consume 24 GPU-hours. Multiply those derived totals by the matching live table entry. No completion time for your dataset is established here.
Time representative training steps, then include evaluation, checkpoint saves and setup time in the rental window. Torchtune's memory-optimization documentation, accessed 28 September 2026, explains that checkpointing recomputes activations and offloading moves them to CPU. Lower peak VRAM therefore does not mean lower total cost; compare measured completion times for the same quality target.
Start with QLoRA if memory blocks the model you need. Use LoRA when its measured peak fits your budget. Pay for full fine-tuning only after adapters miss your quality target and the improvement justifies the cluster bill.
Sources
- Hu et al., LoRA authors and submission, 17 June 2021; mechanism and reported reductions, 16 October 2021.
- Dettmers et al., QLoRA abstract and techniques and experiments, 23 May 2023.
- Hugging Face, model memory anatomy, accessed 28 September 2026.
- NVIDIA, model weight storage guidance, 1 September 2026.
- ZeRO paper, 13 May 2020.
- Torchtune, QLoRA tutorial and Llama2-7B recipe, accessed 28 September 2026.
- Hugging Face PEFT, published memory comparison, accessed 28 September 2026.
- Biderman et al., LoRA Learns Less and Forgets Less, revised 20 September 2024.
- Hugging Face PEFT, LoRA API and quantization guide, accessed 28 September 2026.
- Hugging Face TRL, Axolotl, torchtune overview and LLaMA-Factory, accessed 28 September 2026.
- Unsloth, fine-tuning guide, accessed 28 September 2026.
- Torchtune, memory optimizations, accessed 28 September 2026.