Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.

Selecting the Right GPU for Qwen3 Inference

A practical playbook for LLM inference on GPUs: roofline thinking, bandwidth vs. compute, precision and utilization by model size.

6 min readflozi00
aidatacentergpudeep-learninghardwareguide

Which GPU fits Qwen3 inference?

First check that weights, KV cache and runtime allocations fit in GPU memory; then compare bandwidth and compute for your batch size and context. This guide applies that calculation to RTX PRO 6000, H200 and DGX Station. Its token-rate examples are theoretical ceilings, not measured serving performance.

Use the VRAM calculator for a memory estimate and the GPU buying guide for specifications and purchasing criteria. This playbook explains the hardware-model arithmetic; the vLLM vs SGLang guide covers serving-engine selection.

This overview extends the calculations from LLM Inference Math: From Theory to Hardware and applies them to concrete hardware + model pairings. The goal is to make it obvious when the NVIDIA RTX PRO 6000, H200, or the shipping DGX Station delivers the best efficiency for Qwen3-class workloads ranging from 4B to 32B active parameters.

Important: Every number in this playbook is a theoretical ceiling derived from vendor specs and simplified roofline math. Real deployments often land lower because kernels are imperfect, host↔device pipelines add friction, and GPUs rarely sustain 100% efficiency across an entire decode pass.

Recap: the 60-second ops:byte checklist

  1. Compute the GPU's ops:byte ratio (peak FLOPS ÷ memory bandwidth).
  2. Estimate the arithmetic intensity of the whole operation as FLOPs divided by bytes moved. Attention-only d_head ÷ 2 cannot classify the full transformer pass.
  3. Compare that intensity with the GPU ratio. Below the ratio, the simplified roofline is memory-bound; above it, compute can bind.
  4. For a small-batch dense decode, weight_bytes ÷ memory_bandwidth_bytes_per_second is a weight-read time floor per step, provided weights and KV fit. KV reads and other work increase time.

GPU capability snapshots

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

  • 96 GB GDDR7 ECC VRAM fed by 1,792 GB/s bandwidth and 503.8 TFLOPS FP16/BF16 Tensor (dense) compute — the ratio that matters for inference — which yields an ops:byte ratio near 281. (The ~125 TFLOPS non-tensor FP32 shader rate would misleadingly suggest ~70; see the footnote in the inference math guide.)
  • Peak board power of 600 W enables deskside deployments where noise + thermals matter but rack power is limited.
  • Ideal when you need local fine-tuning or multi-modal prototyping (vision, audio) with up to ~64 GB of weights plus a useful KV cache budget.1

NVIDIA H200 Tensor Core GPU (SXM + NVL)

  • First Hopper-based accelerator with 141 GB of HBM3e and 4.8 TB/s of memory bandwidth; BF16 Tensor performance reaches 989 TFLOPS dense (1,979 TFLOPS with sparsity), so ops:byte climbs to roughly 206 (dense) / 412 (sparse).
  • Ships with hardware MIG slicing (7 instances) and optional NVL configurations for air-cooled racks.
  • Best suited for 14B+ dense models or MoE deployments where both capacity and streaming bandwidth dominate cost.2

NVIDIA DGX Station (Grace Blackwell Ultra)

  • Desktop system combining one Blackwell-Ultra GPU (252 GB HBM3e @ 7.1 TB/s) with a 72-core Grace CPU and 496 GB LPDDR5X (396 GB/s). The combined 748 GB is coherent addressable memory, but only 252 GB has HBM bandwidth.3
  • NVLink-C2C connects CPU and GPU at 900 GB/s aggregate. Data in Grace memory is still limited by its own memory bandwidth and access pattern; account for this when offloading weights or indices.
  • The datasheet rates the GPU at 20 PFLOPS FP4 / 5 PFLOPS BF16 (sparse; 15 / 2.5 dense) — the same Blackwell Ultra silicon as HGX B300 servers, whose 8-GPU baseboard NVIDIA rates at 18 PFLOPS dense BF16 (~2.25 PFLOPS per GPU).34
  • Targets multi-user labs that need on-prem autonomy for iterative training, MoE routing experiments, and agent stacks before shipping them to a cluster.

Qwen3 model footprints (batch size 1, BF16)

Each model reuses the GQA-aware KV-cache formula (2 * layers * num_kv_heads * head_dim * 2 bytes; all Qwen3 models below use 8 KV heads with head_dim 128).

ModelActive paramsHidden size / layersWeights (GB)KV cache per token (MB)Notes
Qwen3-4B-Instruct~4B2,560 / 36~80.15Sliding-window ready; great for CPU offload experiments.5
Qwen3-VL-8B-Instruct~8B4,096 / 36~160.15Vision-language encoder adds ~1152-dim vision tower.6
Qwen3-14B~14B5,120 / 40~280.1640-layer stack with 1M rope theta for 40k context.7
Qwen3-32B~32.8B5,120 / 64~65.50.27Dense 64-layer decoder; 0.27 GB is KV per 1k tokens at bf16.8

Matching scenarios

1. Workstation prototyping (RTX PRO 6000)

  • Recommended models: Qwen3-4B, Qwen3-VL-8B.
  • Why: Approximate bf16 weight payloads of 8–16 GB leave substantial space on a 96 GB card, but vision weights, KV cache and runtime allocations must be included for the actual checkpoint.
  • Throughput: 1.792 TB/s ÷ weight bytes gives weight-only batch-1 ceilings of ~224 tok/s for 8 GB or ~112 tok/s for 16 GB. They are not expected serving rates.
  • Tip: Test larger batches against the latency target and available KV memory; sidecar processes consume memory and compute too.

2. Enterprise copilots (H200 SXM)

  • Recommended models: Qwen3-14B and Qwen3-32B, both dense.
  • Why: 141 GB HBM3e holds ~65.5 GB of Qwen3-32B bf16 weights, leaving ~75.5 GB before KV and runtime reservations. Weight-only batch-1 ceilings are ~170 tok/s for an approximate 28 GB 14B checkpoint and ~73 tok/s for Qwen3-32B at 4.8 TB/s. Actual rates may be lower and the binding limit can change with batch and context.
  • Tip: MIG partitions the GPU's capacity and compute among tenants; choose instance sizes that fit each model and measure the resulting rates.

3. Lab-scale supercomputer (DGX Station)

  • Recommended models: Qwen3 variants and colocated tools whose combined active GPU working set fits within the 252 GB HBM budget.
  • Why: Grace's 496 GB LPDDR5X can hold CPU-side datasets, but weights or KV spilled there do not run at the 7.1 TB/s HBM bandwidth. Account separately for both memory tiers and the 900 GB/s C2C link.
  • Throughput: For Qwen3-32B's ~65.5 GB bf16 weights, 7.1 TB/s ÷ 65.5 GB ≈ 108 tok/s is a weight-read ceiling at batch 1, before KV, communication or compute costs.4
  • Tip: MIG supports partitioning into up to seven instances, but reserved slices remove resources from the main model; test whether partitioning helps your workload.

Quick pairing matrix

ScenarioModelGPUEst. tokens/s (batch 1)Primary bottleneckNotes
Edge copilotsQwen3-4BRTX PRO 6000~220Memory BWPlenty of VRAM left for RAG embeddings.
Vision agent demosQwen3-VL-8BRTX PRO 6000~110Memory BWVision tower benefits from 96 GB VRAM for image batches.
Customer support copilotsQwen3-14BH200~170Memory BWMIG lets you mirror-prod topology in dev.
Technical assistant / codegenQwen3-32BH200~73Weight-read floorA single H200 can hold several ~2k-context sequences; actual batch limit depends on KV and runtime memory.
Multi-agent sandboxQwen3-32B + toolsDGX Station~108Weight-read floorKeep GPU-active weights and KV within 252 GB HBM; Grace RAM is a slower tier.

Interactive calculator

Use the LLM inference planner to stress-test context windows, batch sizes, and precision assumptions against each GPU profile — it computes VRAM fit, decode throughput, prefill latency and the break-even batch B* live for every model and GPU in this playbook.

Note: Tokens per second plateau once the workload hits the compute roofline—the calculator compares both limits and reports the stricter one.

Footnotes

  1. NVIDIA RTX PRO 6000 Blackwell Workstation Edition specifications, NVIDIA. ↩

  2. NVIDIA H200 Tensor Core GPU specifications, NVIDIA. ↩

  3. NVIDIA DGX Station (Grace Blackwell Ultra) specifications, NVIDIA. ↩ ↩2

  4. NVIDIA HGX Platform and Blackwell Ultra specifications (HGX B300), NVIDIA. ↩ ↩2

  5. Qwen/Qwen3-4B-Instruct-2507 model card, Hugging Face. ↩

  6. Qwen/Qwen3-VL-8B-Instruct model card, Hugging Face. ↩

  7. Qwen/Qwen3-14B model card, Hugging Face. ↩

  8. Qwen/Qwen3-32B model card, Hugging Face. ↩