Skip to content
Friendly disclaimer: flozi00 TechHub is a solo side-project next to a full-time job — personal learning notes, no official statements. Verify critical steps yourself.
← Home

Find the guide for your question

Start with the question you want to answer. Each reading path explains what its guides cover, so you can choose a foundation, calculator or deeper analysis.

How much VRAM does my LLM need?

Start with the memory budget, then calculate weights, context and runtime overhead. Fine-tuning has a separate budget.

  1. LLM VRAM Requirements: Weights, Quantization & KV Cache

    Understand the inference memory budget and its assumptions.

  2. LLM VRAM Calculator: GPU Memory & Inference Estimates

    Calculate model fit for a chosen GPU, precision and context.

  3. LLM Quantization: FP8, INT4, GGUF & VRAM Requirements

    Separate weight compression from quality and runtime constraints.

  4. KV Cache Explained: Memory Formula, GQA, MQA & MLA

    Work out the memory consumed by context and concurrent requests.

  5. LoRA and QLoRA Fine-Tuning: The Actual Memory Math

    Plan adapters, optimizer states and activations for fine-tuning.

Which GPU and interconnect fit my workload?

Use model fit and bandwidth for inference planning. Read training benchmarks as system results and distinguish them from purchasing evidence.

  1. GPU Buying Guide for LLM Inference (2026)

    Compare specifications, quotes and workload requirements.

  2. GPU for Qwen3 Inference: VRAM, Bandwidth & Model Fit

    Apply roofline math to concrete Qwen3 and GPU pairings.

  3. NVIDIA B200 vs GB200: Differences & MLPerf Benchmarks

    Compare published B200 and GB200 training submissions fairly.

  4. NVIDIA NVLink: Bandwidth, NVSwitch & GPU Topology

    Understand what an NVLink domain connects.

  5. Scale-Up vs Scale-Out: NVLink and All-Reduce

    Calculate communication within and between GPU servers.

  6. AMD Instinct vs NVIDIA: Memory and Serving Limits

    Compare AMD and NVIDIA memory specifications and their limits.

How do I choose and operate an LLM serving engine?

Keep local deployment, engine selection, release maintenance and cost calculations separate. Each guide answers a different operational question.

  1. llama.cpp vs Ollama vs vLLM: Local LLM Serving Compared

    Choose between local runtimes and GPU serving.

  2. vLLM vs SGLang: Serving, Caching & When to Use Each

    Compare vLLM and SGLang for your request patterns.

  3. LLM Serving Engines 2026: Upgrades and Version Pinning

    Budget for upgrades, version pinning and rollback.

  4. LLM Inference Performance: Roofline Math & GPU Bandwidth

    Understand the formulas behind decode and prefill limits.

  5. LLM Inference Economics: Costs & Cloud Break-Even

    Compare self-hosting costs with an API under explicit assumptions.

How do I size and manage a KV cache?

Move from per-token memory to working sets and prefix reuse. Sizing, eviction and offload solve different problems.

  1. KV Cache Explained: Memory Formula, GQA, MQA & MLA

    Learn the per-token formula and attention-layout differences.

  2. KV Cache Memory Growth Explorer

    Explore how growing context consumes free GPU memory.

  3. KV-Cache Glossary: GQA, MLA, TTFT and Eviction

    Look up serving terms and their precise definitions.

  4. KVSET: Size an LLM Prefix Cache with LRU Traces

    Derive a cache-capacity curve from an LRU trace.

  5. LLM Prefix Caching: LRU Eviction and Its Limits

    Evaluate LRU eviction for different workloads and capacities.

  6. Agent Inference: KV-Cache Traffic and Memory Tiers

    Calculate agent working sets and memory-tier traffic.

What changes between LLM architectures?

Learn the basic token-to-output flow, then compare individual released checkpoints. Architecture guides describe mechanisms and serving limits.

  1. How LLMs Actually Work — An Animated Walkthrough

    Follow embeddings, attention and sampling in an animated walkthrough.

  2. Qwen3-Next-80B-A3B: Architecture and Serving Limits

    Inspect Qwen3-Next hybrid attention and sparse experts.

  3. Kimi K3 Architecture: KDA, MLA and LatentMoE

    Read Kimi K3 recurrence, MLA and LatentMoE routing.

  4. Inside GLM-5.3-Flash: KDA, Sparse MLA, mHC, and 8-of-288 MoE

    Inspect GLM-5.3-Flash attention, residuals and cache shapes.

  5. Cloudflare Clef: Qwen Backbones and Schema Scoring

    Understand Clef decision scoring and its prefill-only design.

How do I get NVIDIA GPUs working on Ubuntu?

Separate installation and container verification from CUDA error diagnosis. A reported workaround is not a universal fix.

  1. NVIDIA GPU Containers on Ubuntu 24.04

    Install the host driver and verify GPU access inside Docker.

  2. NVIDIA Error 802 on Ubuntu: Diagnose Before Trying nokaslr

    Diagnose CUDA error 802 before considering nokaslr.

Browse all technical articles →