Find the guide for your question
Start with the question you want to answer. Each reading path explains what its guides cover, so you can choose a foundation, calculator or deeper analysis.
How much VRAM does my LLM need?
Start with the memory budget, then calculate weights, context and runtime overhead. Fine-tuning has a separate budget.
- LLM VRAM Requirements: Weights, Quantization & KV Cache
Understand the inference memory budget and its assumptions.
- LLM VRAM Calculator: GPU Memory & Inference Estimates
Calculate model fit for a chosen GPU, precision and context.
- LLM Quantization: FP8, INT4, GGUF & VRAM Requirements
Separate weight compression from quality and runtime constraints.
- KV Cache Explained: Memory Formula, GQA, MQA & MLA
Work out the memory consumed by context and concurrent requests.
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
Plan adapters, optimizer states and activations for fine-tuning.
Which GPU and interconnect fit my workload?
Use model fit and bandwidth for inference planning. Read training benchmarks as system results and distinguish them from purchasing evidence.
- GPU Buying Guide for LLM Inference (2026)
Compare specifications, quotes and workload requirements.
- GPU for Qwen3 Inference: VRAM, Bandwidth & Model Fit
Apply roofline math to concrete Qwen3 and GPU pairings.
- NVIDIA B200 vs GB200: Differences & MLPerf Benchmarks
Compare published B200 and GB200 training submissions fairly.
- NVIDIA NVLink: Bandwidth, NVSwitch & GPU Topology
Understand what an NVLink domain connects.
- Scale-Up vs Scale-Out: NVLink and All-Reduce
Calculate communication within and between GPU servers.
- AMD Instinct vs NVIDIA: Memory and Serving Limits
Compare AMD and NVIDIA memory specifications and their limits.
How do I choose and operate an LLM serving engine?
Keep local deployment, engine selection, release maintenance and cost calculations separate. Each guide answers a different operational question.
- llama.cpp vs Ollama vs vLLM: Local LLM Serving Compared
Choose between local runtimes and GPU serving.
- vLLM vs SGLang: Serving, Caching & When to Use Each
Compare vLLM and SGLang for your request patterns.
- LLM Serving Engines 2026: Upgrades and Version Pinning
Budget for upgrades, version pinning and rollback.
- LLM Inference Performance: Roofline Math & GPU Bandwidth
Understand the formulas behind decode and prefill limits.
- LLM Inference Economics: Costs & Cloud Break-Even
Compare self-hosting costs with an API under explicit assumptions.
How do I size and manage a KV cache?
Move from per-token memory to working sets and prefix reuse. Sizing, eviction and offload solve different problems.
- KV Cache Explained: Memory Formula, GQA, MQA & MLA
Learn the per-token formula and attention-layout differences.
- KV Cache Memory Growth Explorer
Explore how growing context consumes free GPU memory.
- KV-Cache Glossary: GQA, MLA, TTFT and Eviction
Look up serving terms and their precise definitions.
- KVSET: Size an LLM Prefix Cache with LRU Traces
Derive a cache-capacity curve from an LRU trace.
- LLM Prefix Caching: LRU Eviction and Its Limits
Evaluate LRU eviction for different workloads and capacities.
- Agent Inference: KV-Cache Traffic and Memory Tiers
Calculate agent working sets and memory-tier traffic.
What changes between LLM architectures?
Learn the basic token-to-output flow, then compare individual released checkpoints. Architecture guides describe mechanisms and serving limits.
- How LLMs Actually Work — An Animated Walkthrough
Follow embeddings, attention and sampling in an animated walkthrough.
- Qwen3-Next-80B-A3B: Architecture and Serving Limits
Inspect Qwen3-Next hybrid attention and sparse experts.
- Kimi K3 Architecture: KDA, MLA and LatentMoE
Read Kimi K3 recurrence, MLA and LatentMoE routing.
- Inside GLM-5.3-Flash: KDA, Sparse MLA, mHC, and 8-of-288 MoE
Inspect GLM-5.3-Flash attention, residuals and cache shapes.
- Cloudflare Clef: Qwen Backbones and Schema Scoring
Understand Clef decision scoring and its prefill-only design.
How do I get NVIDIA GPUs working on Ubuntu?
Separate installation and container verification from CUDA error diagnosis. A reported workaround is not a universal fix.
- NVIDIA GPU Containers on Ubuntu 24.04
Install the host driver and verify GPU access inside Docker.
- NVIDIA Error 802 on Ubuntu: Diagnose Before Trying nokaslr
Diagnose CUDA error 802 before considering nokaslr.