Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job โ€” personal learning notes, no official statements. Verify critical steps yourself.

LLM VRAM Requirements: A Mathematical Deep Dive

Qwen3-VL-32B GPU-memory arithmetic for weights, KV cache and inference, with separate model-state accounting for full training.

9 min readflozi00
aimachine-learninggpudatacenterdeep-learning

Deploying Large Language Models (LLMs) requires careful consideration of GPU memory. This guide works through an illustrative inference estimate for Qwen3-VL-32B-Instruct and separately accounts for full-training model states.

How much VRAM does an LLM need?

For inference, budget model weights + KV cache + runtime allocations. The weight payload alone is approximately parameters ร— bits per weight รท 8 bytes. A dense 32B model therefore needs about 64 GB at BF16, 32 GB at 8 bits or 16 GB at 4 bits just for weights, before quantization scales, KV cache and runtime overhead. These are decimal GB estimates, not a guarantee that a model will fit on a card of that size.

Context length and concurrent sequences increase KV-cache memory. This guide calculates those components separately for Qwen3-VL-32B, then distinguishes inference from full training. To explore a model, GPU and context configuration directly, open the LLM VRAM calculator. For format overhead, see the quantization guide; for cache formulas, see KV cache explained.

Why VRAM Calculation Matters

Before deploying an LLM, you need to answer critical questions:

  • How much GPU memory will my model consume?
  • Can I run this model on my current hardware?
  • Which quantization method provides the best memory-performance tradeoff?
  • How many concurrent users can I support?

This guide provides the mathematical foundation to answer these questions accurately.

Example Model: Qwen3-VL-32B-Instruct

We'll use Qwen3-VL-32B-Instruct as our reference model throughout this guide. This multimodal model combines vision and language capabilities with the following architecture:

ParameterValueDescription
Model Parameters33.4 billionTotal parameters (33,357,390,064 per HF safetensors)
Hidden Size5,120Dimension of hidden representations
Intermediate Size25,600FFN intermediate dimension (5x hidden size)
Number of Layers64Total transformer blocks
Attention Heads64Number of query attention heads
KV Heads8Number of key-value heads (GQA)
Head Dimension128Dimension per attention head
Max Context Length262,144Maximum sequence length (256k tokens)
ArchitectureGrouped Query AttentionUses GQA for efficient inference

Configuration Source: The model configuration is extracted from the text_config section of the model's config.json file on Hugging Face Hub.

Core Memory Components

VRAM consumption for LLMs consists of four primary components:

1. Model Weights Memory

The base memory required to store the model's parameters.

Model Weights (bytes) = Number of Parameters ร— Bytes per Parameter

Bytes per Parameter depends on the data type (quantization level):

Data TypeBytes per ParameterPrecision
float324 bytesFull precision
float16/bfloat162 bytesHalf precision
int8/fp81 byte8-bit quantization
int4/fp40.5 bytes4-bit quantization

Example Calculation for Qwen3-VL-32B:

Number of Parameters: 33,357,390,064 (~33.4B)

float32:  33,357,390,064 ร— 4.0   = 133,429,560,256 bytes = 133.43 GB
float16:  33,357,390,064 ร— 2.0   = 66,714,780,128 bytes  = 66.71 GB
int8:     33,357,390,064 ร— 1.0   = 33,357,390,064 bytes  = 33.36 GB
int4:     33,357,390,064 ร— 0.5   = 16,678,695,032 bytes  = 16.68 GB

2. KV Cache Memory

The Key-Value cache stores intermediate attention states for efficient autoregressive generation. This is the most significant dynamic memory component during inference.

KV Cache (bytes) = 2 ร— Batch Size ร— Sequence Length ร— Num Layers ร— Num KV Heads ร— Head Dimension ร— KV Data Type Size

Breaking Down the Formula:

  • 2ร—: Separate storage for Keys and Values
  • Batch Size: Number of concurrent requests
  • Sequence Length: Maximum context length (input + output)
  • Num Layers: Number of transformer blocks
  • Num KV Heads: Number of key-value heads (8 for GQA in Qwen3-VL)
  • Head Dimension: Size of each attention head (128)
  • KV Data Type Size: Bytes per value (typically 2 for float16)

Example Calculation for Qwen3-VL-32B:

Scenario: 1 user, 8,192 token context, float16 KV cache

Batch Size: 1
Sequence Length: 8,192 tokens
Num Layers: 64
Num KV Heads: 8 (Grouped Query Attention)
Head Dimension: 128
KV Data Type: float16 (2 bytes)

KV Cache = 2 ร— 1 ร— 8,192 ร— 64 ร— 8 ร— 128 ร— 2
         = 2 ร— 1 ร— 8,192 ร— 64 ร— 8 ร— 256
         = 2 ร— 1,073,741,824 bytes
         = 2,147,483,648 bytes
         = 2.15 GB (decimal; = 2.00 GiB binary)

Scaling with Batch Size:

Batch SizeUsersKV Cache Memory (float16)
11 concurrent user2.15 GB
44 concurrent users8.59 GB
88 concurrent users17.18 GB
1616 concurrent users34.36 GB

Scaling with Sequence Length:

Sequence LengthContext SizeKV Cache Memory (batch=1, float16)
2,0482k tokens0.54 GB
8,1928k tokens2.15 GB
32,76832k tokens8.59 GB
131,072128k tokens34.36 GB

3. Activation Memory

Memory required for intermediate computations during forward passes. The formula below is a planning heuristic, not an exact PyTorch or vLLM allocator formula; attention implementation, prefill chunking, batching and temporary buffers change peak memory.

Illustrative Activation Reserve = Batch Size ร— Sequence Length ร— (18 ร— Hidden Size + 4 ร— Intermediate Size)

Example Calculation for Qwen3-VL-32B:

Scenario: 1 user, 8,192 token context

Batch Size: 1
Sequence Length: 8,192
Hidden Size: 5,120
Intermediate Size: 25,600

Activation Memory = 1 ร— 8,192 ร— (18 ร— 5,120 + 4 ร— 25,600)
                  = 8,192 ร— (92,160 + 102,400)
                  = 8,192 ร— 194,560
                  = 1,593,835,520 bytes
                  = 1.59 GB (base value)

Data Type Multipliers:

Different quantization levels have different activation memory footprints:

Data TypeMultiplierEffective Activation Memory
float322.0ร—1.59 ร— 2.0 = 3.19 GB
float16/bfloat161.0ร—1.59 ร— 1.0 = 1.59 GB
int8/fp81.0ร—1.59 ร— 1.0 = 1.59 GB
int4/fp4 (weight-only)1.0ร—1.59 ร— 1.0 = 1.59 GB

4. Non-PyTorch Memory Overhead

System-level memory overhead for CUDA context, cuBLAS, and other framework components.

Illustrative Runtime Reserve โ‰ˆ 1.00 GB

The 1 GB is an illustrative reserve, not a fixed framework cost. The interactive calculator instead uses 1.2 GB scaled by its format-overhead factor (for example ~1.26 GB for int8). Measure the actual runtime peak before choosing a device.

Complete Inference Memory Formula

Combining all components, the total VRAM required for inference:

Total Inference VRAM = (Model Weights + KV Cache + Non-PyTorch Memory + Activations) / GPU Utilization

Capacity reserve: The examples divide by 0.9 to set aside 10% of VRAM. This is a planning choice, not an observed runtime utilization.

Calculator difference: These hand examples use raw payload bytes and an explicit 10% capacity reserve. The interactive calculator adds assumed weight-format overhead โ€” FP8 ร—1.03, INT8 ร—1.05, INT4/AWQ ร—1.12; its NVFP4 entry uses 4.5 effective bits per weight, including block scales. It uses a 1.2 GB runtime reserve multiplied by the format factor and does not divide capacity by the throughput-efficiency slider. The two estimates answer slightly different planning questions; neither guarantees fit.

Example: Qwen3-VL-32B Inference (int8 quantization)

Configuration:

  • Quantization: int8 (1 byte per parameter)
  • Batch Size: 1 user
  • Sequence Length: 8,192 tokens
  • KV Cache Data Type: float16
  • GPU Utilization: 0.9

Step-by-Step Calculation:

1. Model Weights:
   33,357,390,064 ร— 1 byte = 33,357,390,064 bytes = 33.36 GB

2. KV Cache (float16):
   2 ร— 1 ร— 8,192 ร— 64 ร— 8 ร— 128 ร— 2 = 2,147,483,648 bytes = 2.15 GB

3. Activations (int8 uses 1.0ร— multiplier):
   1 ร— 8,192 ร— (18 ร— 5,120 + 4 ร— 25,600) ร— 1.0 = 1,593,835,520 bytes = 1.59 GB

4. Non-PyTorch Memory:
   โ‰ˆ 1.00 GB

5. Total (before GPU utilization adjustment):
   33.36 + 2.15 + 1.59 + 1.00 = 38.10 GB

6. Adjusted for GPU Utilization (90%):
   38.10 / 0.9 = 42.33 GB

Result: You need approximately 42 GB of VRAM to run Qwen3-VL-32B in int8 quantization with 8k context for a single user.

Example: Qwen3-VL-32B Inference (int4 quantization)

Configuration:

  • Quantization: int4 (0.5 bytes per parameter)
  • Batch Size: 4 users
  • Sequence Length: 8,192 tokens
  • KV Cache Data Type: float16
  • GPU Utilization: 0.9

Step-by-Step Calculation:

1. Model Weights:
   33,357,390,064 ร— 0.5 bytes = 16,678,695,032 bytes = 16.68 GB

2. KV Cache (float16, batch=4):
   2 ร— 4 ร— 8,192 ร— 64 ร— 8 ร— 128 ร— 2 = 8,589,934,592 bytes = 8.59 GB

3. Activations (int4 uses 1.0ร— multiplier, batch=4):
   4 ร— 8,192 ร— (18 ร— 5,120 + 4 ร— 25,600) ร— 1.0 = 6,375,342,080 bytes = 6.38 GB

4. Non-PyTorch Memory:
   โ‰ˆ 1.00 GB

5. Total (before GPU utilization adjustment):
   16.68 + 8.59 + 6.38 + 1.00 = 32.65 GB

6. Adjusted for GPU Utilization (90%):
   32.65 / 0.9 = 36.28 GB

Result: You need approximately 36 GB of VRAM to run Qwen3-VL-32B in int4 quantization with 8k context for 4 concurrent users.

Training Memory Requirements

Full-parameter training has a different memory budget from inference. There is no persistent inference KV cache to add: automatic differentiation keeps or recomputes forward activations for the backward pass. Activation memory depends on the attention kernel, checkpointing, microbatch size and sequence length.

For 33,357,390,064 trainable parameters, a simple mixed-precision Adam accounting is:

StateBytes per parameterTotal
bf16 weights266.71 GB
bf16 gradients266.71 GB
Adam moments in fp328266.86 GB
Subtotal without fp32 master weights12400.29 GB
Optional fp32 master weights+4+133.43 GB
Subtotal with fp32 master weights16533.72 GB

These are model-state totals across all devices, before activations, temporary buffers, runtime allocations and any safety margin. Optimizer implementation, parameter sharding and offload change where that state resides. Quantized LoRA/QLoRA training freezes the base weights and trains small adapters, so it requires a separate calculation; see LoRA/QLoRA memory math. A device count cannot be obtained by simply dividing the model-state total by a card's VRAM: the chosen parallelism determines how weights, optimizer state and activations are distributed.

GPU Memory Comparison Table

Here's a comprehensive comparison for Qwen3-VL-32B across different quantization methods:

QuantizationRaw weight payloadKV cacheActivation heuristicRuntime reserveInference estimate including 10% reserve
float32133.43 GB2.15 GB3.19 GB1.00 GB155.30 GB
float1666.71 GB2.15 GB1.59 GB1.00 GB79.39 GB
int833.36 GB2.15 GB1.59 GB1.00 GB42.33 GB
int416.68 GB2.15 GB1.59 GB1.00 GB23.80 GB

Configuration: Batch size 1, Sequence length 8,192 tokens, GPU utilization 0.9

Practical GPU Recommendations

Based on the calculations above, here are suitable GPU configurations for Qwen3-VL-32B:

Inference Deployment

QuantizationRequired VRAMRecommended GPUsUse Case
int4 (24 GB)32 GB comfortable1ร— RTX 5090 (32 GB) or RTX PRO 6000 (96 GB)Cost-effective inference
int8 (42 GB)48 GB minimum1ร— L40S (48 GB) or RTX PRO 6000 (96 GB)Higher quality inference
float16 (79 GB)96 GB tight / 141 GB with headroom1ร— RTX PRO 6000 Blackwell (96 GB) or 1ร— H200 (141 GB)Full precision inference

Full training requires a separate system design for optimizer states, sharding, activations and checkpointing; the training section above gives the model-state subtotal.

Key Takeaways

  1. Model Weights Scale Linearly: Doubling parameters doubles weight memory
  2. KV Cache Scales with Context: KV cache memory grows linearly with context length
  3. Batch Size Multiplies KV Cache: Each concurrent user adds KV cache overhead
  4. Quantization Dramatically Reduces Memory: int4 uses ~1/8th the memory of float32
  5. Training memory is method-dependent: Full training, LoRA and QLoRA have different trainable states and activation footprints.
  6. Leave capacity headroom: Choose a reserve for the actual runtime and verify peak allocation on hardware.

Advanced Considerations

Grouped Query Attention (GQA) Impact

Qwen3-VL-32B uses Grouped Query Attention with 8 KV heads instead of 64 query heads. This reduces KV cache memory by 8ร— compared to standard Multi-Head Attention:

Standard MHA: 64 KV heads โ†’ KV Cache = 17.18 GB (8k context, float16)
GQA (Qwen3): 8 KV heads โ†’ KV Cache = 2.15 GB (8k context, float16)

Memory Savings: 15.03 GB (87.5% reduction)

Long Context Scaling

Qwen3-VL-32B supports up to 262k tokens. Here's how KV cache scales:

Context LengthKV Cache (float16, batch=1)Recommended VRAM
8k tokens2.15 GB48 GB GPU
32k tokens8.59 GB64 GB GPU
128k tokens34.36 GB96 GB GPU
256k tokens68.72 GB128 GB GPU

KV Cache Quantization

Using fp8 for KV cache can halve memory usage:

float16 KV Cache (8k, batch=1): 2.15 GB
fp8 KV Cache (8k, batch=1): 1.07 GB

Savings: 50% memory reduction with minimal quality loss

Conclusion

Accurate VRAM calculation is essential for efficient LLM deployment. By understanding the mathematical foundations:

  1. Choose the right quantization for your quality-memory tradeoff
  2. Plan for KV cache scaling with concurrent users and context length
  3. Budget for training overhead if fine-tuning
  4. Select appropriate GPUs based on actual requirements
  5. Always include safety margins to prevent OOM errors

Use these formulas to estimate memory requirements for any Hugging Face model by extracting configuration parameters and applying the calculations shown in this guide.

References and Tools

  • Model Configuration: Qwen3-VL-32B-Instruct config.json
  • Hugging Face Hub: Model metadata and parameter counts available via API
  • GPU Specifications: Check manufacturer specifications for exact VRAM capacities
  • Quantization Methods: GPTQ, AWQ, GGUF for different precision levels