Which GPU fits Qwen3 inference?
First check that weights, KV cache and runtime allocations fit in GPU memory; then compare bandwidth and compute for your batch size and context. This guide applies that calculation to RTX PRO 6000, H200 and DGX Station. Its token-rate examples are theoretical ceilings, not measured serving performance.
Use the VRAM calculator for a memory estimate and the GPU buying guide for specifications and purchasing criteria. This playbook explains the hardware-model arithmetic; the vLLM vs SGLang guide covers serving-engine selection.
This overview extends the calculations from LLM Inference Math: From Theory to Hardware and applies them to concrete hardware + model pairings. The goal is to make it obvious when the NVIDIA RTX PRO 6000, H200, or the shipping DGX Station delivers the best efficiency for Qwen3-class workloads ranging from 4B to 32B active parameters.
Important: Every number in this playbook is a theoretical ceiling derived from vendor specs and simplified roofline math. Real deployments often land lower because kernels are imperfect, host↔device pipelines add friction, and GPUs rarely sustain 100% efficiency across an entire decode pass.
Recap: the 60-second ops:byte checklist
- Compute the GPU's ops:byte ratio (
peak FLOPS ÷ memory bandwidth). - Estimate the arithmetic intensity of the whole operation as FLOPs divided by bytes moved. Attention-only
d_head ÷ 2cannot classify the full transformer pass. - Compare that intensity with the GPU ratio. Below the ratio, the simplified roofline is memory-bound; above it, compute can bind.
- For a small-batch dense decode,
weight_bytes ÷ memory_bandwidth_bytes_per_secondis a weight-read time floor per step, provided weights and KV fit. KV reads and other work increase time.
GPU capability snapshots
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
- 96 GB GDDR7 ECC VRAM fed by 1,792 GB/s bandwidth and 503.8 TFLOPS FP16/BF16 Tensor (dense) compute — the ratio that matters for inference — which yields an ops:byte ratio near 281. (The ~125 TFLOPS non-tensor FP32 shader rate would misleadingly suggest ~70; see the footnote in the inference math guide.)
- Peak board power of 600 W enables deskside deployments where noise + thermals matter but rack power is limited.
- Ideal when you need local fine-tuning or multi-modal prototyping (vision, audio) with up to ~64 GB of weights plus a useful KV cache budget.1
NVIDIA H200 Tensor Core GPU (SXM + NVL)
- First Hopper-based accelerator with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth; BF16 Tensor performance reaches 989 TFLOPS dense (1,979 TFLOPS with sparsity), so ops:byte climbs to roughly 206 (dense) / 412 (sparse).
- Ships with hardware MIG slicing (7 instances) and optional NVL configurations for air-cooled racks.
- Best suited for 14B+ dense models or MoE deployments where both capacity and streaming bandwidth dominate cost.2
NVIDIA DGX Station (Grace Blackwell Ultra)
- Desktop system combining one Blackwell-Ultra GPU (252 GB HBM3e @ 7.1 TB/s) with a 72-core Grace CPU and 496 GB LPDDR5X (396 GB/s). The combined 748 GB is coherent addressable memory, but only 252 GB has HBM bandwidth.3
- NVLink-C2C connects CPU and GPU at 900 GB/s aggregate. Data in Grace memory is still limited by its own memory bandwidth and access pattern; account for this when offloading weights or indices.
- The datasheet rates the GPU at 20 PFLOPS FP4 / 5 PFLOPS BF16 (sparse; 15 / 2.5 dense) — the same Blackwell Ultra silicon as HGX B300 servers, whose 8-GPU baseboard NVIDIA rates at 18 PFLOPS dense BF16 (~2.25 PFLOPS per GPU).34
- Targets multi-user labs that need on-prem autonomy for iterative training, MoE routing experiments, and agent stacks before shipping them to a cluster.
Qwen3 model footprints (batch size 1, BF16)
Each model reuses the GQA-aware KV-cache formula (2 * layers * num_kv_heads * head_dim * 2 bytes; all Qwen3 models below use 8 KV heads with head_dim 128).
| Model | Active params | Hidden size / layers | Weights (GB) | KV cache per token (MB) | Notes |
|---|---|---|---|---|---|
| Qwen3-4B-Instruct | ~4B | 2,560 / 36 | ~8 | 0.15 | Sliding-window ready; great for CPU offload experiments.5 |
| Qwen3-VL-8B-Instruct | ~8B | 4,096 / 36 | ~16 | 0.15 | Vision-language encoder adds ~1152-dim vision tower.6 |
| Qwen3-14B | ~14B | 5,120 / 40 | ~28 | 0.16 | 40-layer stack with 1M rope theta for 40k context.7 |
| Qwen3-32B | ~32.8B | 5,120 / 64 | ~65.5 | 0.27 | Dense 64-layer decoder; 0.27 GB is KV per 1k tokens at bf16.8 |
Matching scenarios
1. Workstation prototyping (RTX PRO 6000)
- Recommended models: Qwen3-4B, Qwen3-VL-8B.
- Why: Approximate bf16 weight payloads of 8–16 GB leave substantial space on a 96 GB card, but vision weights, KV cache and runtime allocations must be included for the actual checkpoint.
- Throughput:
1.792 TB/s ÷ weight bytesgives weight-only batch-1 ceilings of ~224 tok/s for 8 GB or ~112 tok/s for 16 GB. They are not expected serving rates. - Tip: Test larger batches against the latency target and available KV memory; sidecar processes consume memory and compute too.
2. Enterprise copilots (H200 SXM)
- Recommended models: Qwen3-14B and Qwen3-32B, both dense.
- Why: 141 GB HBM3e holds ~65.5 GB of Qwen3-32B bf16 weights, leaving ~75.5 GB before KV and runtime reservations. Weight-only batch-1 ceilings are ~170 tok/s for an approximate 28 GB 14B checkpoint and ~73 tok/s for Qwen3-32B at 4.8 TB/s. Actual rates may be lower and the binding limit can change with batch and context.
- Tip: MIG partitions the GPU's capacity and compute among tenants; choose instance sizes that fit each model and measure the resulting rates.
3. Lab-scale supercomputer (DGX Station)
- Recommended models: Qwen3 variants and colocated tools whose combined active GPU working set fits within the 252 GB HBM budget.
- Why: Grace's 496 GB LPDDR5X can hold CPU-side datasets, but weights or KV spilled there do not run at the 7.1 TB/s HBM bandwidth. Account separately for both memory tiers and the 900 GB/s C2C link.
- Throughput: For Qwen3-32B's ~65.5 GB bf16 weights,
7.1 TB/s ÷ 65.5 GB ≈ 108 tok/sis a weight-read ceiling at batch 1, before KV, communication or compute costs.4 - Tip: MIG supports partitioning into up to seven instances, but reserved slices remove resources from the main model; test whether partitioning helps your workload.
Quick pairing matrix
| Scenario | Model | GPU | Est. tokens/s (batch 1) | Primary bottleneck | Notes |
|---|---|---|---|---|---|
| Edge copilots | Qwen3-4B | RTX PRO 6000 | ~220 | Memory BW | Plenty of VRAM left for RAG embeddings. |
| Vision agent demos | Qwen3-VL-8B | RTX PRO 6000 | ~110 | Memory BW | Vision tower benefits from 96 GB VRAM for image batches. |
| Customer support copilots | Qwen3-14B | H200 | ~170 | Memory BW | MIG lets you mirror-prod topology in dev. |
| Technical assistant / codegen | Qwen3-32B | H200 | ~73 | Weight-read floor | A single H200 can hold several ~2k-context sequences; actual batch limit depends on KV and runtime memory. |
| Multi-agent sandbox | Qwen3-32B + tools | DGX Station | ~108 | Weight-read floor | Keep GPU-active weights and KV within 252 GB HBM; Grace RAM is a slower tier. |
Related
- LLM VRAM Requirements: A Mathematical Deep Dive — the memory math behind every fit call.
- LLM Inference Calculator & Break-Even Analyzer — verify a pairing before you buy it.
Interactive calculator
Use the LLM inference planner to stress-test context windows, batch sizes, and precision assumptions against each GPU profile — it computes VRAM fit, decode throughput, prefill latency and the break-even batch B* live for every model and GPU in this playbook.
Note: Tokens per second plateau once the workload hits the compute roofline—the calculator compares both limits and reports the stricter one.
Footnotes
-
NVIDIA RTX PRO 6000 Blackwell Workstation Edition specifications, NVIDIA. ↩
-
NVIDIA DGX Station (Grace Blackwell Ultra) specifications, NVIDIA. ↩ ↩2
-
NVIDIA HGX Platform and Blackwell Ultra specifications (HGX B300), NVIDIA. ↩ ↩2
-
Qwen/Qwen3-4B-Instruct-2507 model card, Hugging Face. ↩
-
Qwen/Qwen3-VL-8B-Instruct model card, Hugging Face. ↩
-
Qwen/Qwen3-14B model card, Hugging Face. ↩
-
Qwen/Qwen3-32B model card, Hugging Face. ↩