Topics & Tags
All TechHub articles, grouped by tag. Click through the topics.
ai61 articles
- About Florian Zimmermeister
- Qwen3-Next-80B-A3B: Architecture and Serving Limits
- TypeSafe Jev: What the Public Evidence Supports
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- EU AI Act Article 50: Marking and Its Technical Limits
- EU AI Act after the Digital Omnibus: What Moved
- EU AI Act and GDPR for Self-Hosted LLMs
- Parakeet TDT vs Encoder–Decoder ASR: Production Trade-offs
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- CXMT at 10% of DRAM Revenue: What the Figures Show
- Self-hosting DeepSeek V4 Pro: the actual math
- DeepSeek-V4.1-Flash KV-Cache Compression: What the Reported 890 Bytes per Token Includes
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- GPU Buying Guide for LLM Inference (2026)
- Selecting the Right GPU for Qwen3 Inference
- Gumbel Watermarking Finally Shipped — and the Compliance Asymmetry Nobody Priced
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- How LLMs Actually Work — An Animated Walkthrough
- KernelOPT's Agentic Kernel Search: The Verification Cascade Works, The Speedup Fades With Difficulty
- Germany's KI-MIG: Who Enforces the AI Act?
- Inside Kimi K3: Delta Attention, Depth Attention, and LatentMoE
- KREX: Shared-GPU Kernel Benchmarking Without Corrupting the Agent's Search
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- Risk-Controlled KV Eviction: The Reliability Contract and Full-KV Fallback
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
- Leaderboard Margins vs Hidden Model Selection: Which Gains Survive k Secret Variants?
- llama.cpp / Ollama vs vLLM: Which One, When, and Why
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- NVIDIA Agrees to Buy Hugging Face: The Dependency Math Behind the $12.93B Headline
- Greedy Decoding Is Not Precision-Invariant: BF16-vs-FP16 Flips, Margin Gating, and the Kernel-Order Axis
- LLM Prefix Caching: Where LRU Holds Up and Where It Falls Short
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- Quantization Can Change Retrieval Results While Classification Accuracy Holds
- Inside Qwen3.8-Flash-Next: Gated DeltaNet, Sparse Attention, and N-gram Memory
- RAG Retrieval Math: Embedding Memory, Vector Index Size, and the Latency Budget Nobody Calculates
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- The 2026 Serving-Engine Churn Tax: vLLM 0.29, SGLang 0.5.19, and TensorRT-LLM Losing TensorRT
- Snapdragon X2 on Linux: The 80-TOPS NPU Is Not the Story — 228 GB/s Is
- Uncheatable Eval: Scoring LLMs by How Few Bits They Need
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- vLLM vs SGLang: Architecture, Overhead, and When to Pick Which
- My Journey: Obsession, Failure, and AI as an Anchor
- KV Cache Memory Growth Explorer
- LLM Inference Calculator & Break-Even Analyzer
- LLM Inference Economics: Costs & Cloud Break-Even
machine-learning42 articles
- Qwen3-Next-80B-A3B: Architecture and Serving Limits
- TypeSafe Jev: What the Public Evidence Supports
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- Parakeet TDT vs Encoder–Decoder ASR: Production Trade-offs
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- Self-hosting DeepSeek V4 Pro: the actual math
- DeepSeek-V4.1-Flash KV-Cache Compression: What the Reported 890 Bytes per Token Includes
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- Gumbel Watermarking Finally Shipped — and the Compliance Asymmetry Nobody Priced
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- KernelOPT's Agentic Kernel Search: The Verification Cascade Works, The Speedup Fades With Difficulty
- KREX: Shared-GPU Kernel Benchmarking Without Corrupting the Agent's Search
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- Risk-Controlled KV Eviction: The Reliability Contract and Full-KV Fallback
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
- Leaderboard Margins vs Hidden Model Selection: Which Gains Survive k Secret Variants?
- llama.cpp / Ollama vs vLLM: Which One, When, and Why
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- NVIDIA Agrees to Buy Hugging Face: The Dependency Math Behind the $12.93B Headline
- Greedy Decoding Is Not Precision-Invariant: BF16-vs-FP16 Flips, Margin Gating, and the Kernel-Order Axis
- LLM Prefix Caching: Where LRU Holds Up and Where It Falls Short
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- Quantization Can Change Retrieval Results While Classification Accuracy Holds
- RAG Retrieval Math: Embedding Memory, Vector Index Size, and the Latency Budget Nobody Calculates
- The 2026 Serving-Engine Churn Tax: vLLM 0.29, SGLang 0.5.19, and TensorRT-LLM Losing TensorRT
- Uncheatable Eval: Scoring LLMs by How Few Bits They Need
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- vLLM vs SGLang: Architecture, Overhead, and When to Pick Which
gpu33 articles
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- Self-hosting DeepSeek V4 Pro: the actual math
- GPU Buying Guide for LLM Inference (2026)
- Selecting the Right GPU for Qwen3 Inference
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- KernelOPT's Agentic Kernel Search: The Verification Cascade Works, The Speedup Fades With Difficulty
- KREX: Shared-GPU Kernel Benchmarking Without Corrupting the Agent's Search
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- llama.cpp / Ollama vs vLLM: Which One, When, and Why
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- NVIDIA Agrees to Buy Hugging Face: The Dependency Math Behind the $12.93B Headline
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- RAG Retrieval Math: Embedding Memory, Vector Index Size, and the Latency Budget Nobody Calculates
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- The 2026 Serving-Engine Churn Tax: vLLM 0.29, SGLang 0.5.19, and TensorRT-LLM Losing TensorRT
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- vLLM vs SGLang: Architecture, Overhead, and When to Pick Which
- NVIDIA GPU Containers on Ubuntu 24.04
- AMD MI300X/MI350/MI400: More Memory, But Does the Math Win? — A Critical Silicon View of AMD Instinct vs NVIDIA
- NVIDIA B200 vs GB200: What MLPerf Training Actually Shows
- NVIDIA NVLink: Generations, Domains and Verification
- KV Cache Memory Growth Explorer
- LLM Inference Calculator & Break-Even Analyzer
- LLM Inference Economics: Costs & Cloud Break-Even
inference31 articles
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- DeepSeek-V4.1-Flash KV-Cache Compression: What the Reported 890 Bytes per Token Includes
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- Inside Kimi K3: Delta Attention, Depth Attention, and LatentMoE
- KV-Cache Glossary: every term in the serving stack
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- Risk-Controlled KV Eviction: The Reliability Contract and Full-KV Fallback
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- Greedy Decoding Is Not Precision-Invariant: BF16-vs-FP16 Flips, Margin Gating, and the Kernel-Order Axis
- LLM Prefix Caching: Where LRU Holds Up and Where It Falls Short
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- Quantization Can Change Retrieval Results While Classification Accuracy Holds
- Inside Qwen3.8-Flash-Next: Gated DeltaNet, Sparse Attention, and N-gram Memory
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- AMD's MI455X 34× Throughput Claim: What the Disclosure Establishes
- AMD MI300X/MI350/MI400: More Memory, But Does the Math Win? — A Critical Silicon View of AMD Instinct vs NVIDIA
- MLPerf Inference 2026: What a Result Says About Hardware and Software
- KV Cache Memory Growth Explorer
- LLM Inference Calculator & Break-Even Analyzer
- LLM Inference Economics: Costs & Cloud Break-Even
llm25 articles
- Qwen3-Next-80B-A3B: Architecture and Serving Limits
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- DeepSeek-V4.1-Flash KV-Cache Compression: What the Reported 890 Bytes per Token Includes
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- Gumbel Watermarking Finally Shipped — and the Compliance Asymmetry Nobody Priced
- How LLMs Actually Work — An Animated Walkthrough
- KernelOPT's Agentic Kernel Search: The Verification Cascade Works, The Speedup Fades With Difficulty
- Inside Kimi K3: Delta Attention, Depth Attention, and LatentMoE
- KREX: Shared-GPU Kernel Benchmarking Without Corrupting the Agent's Search
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- Risk-Controlled KV Eviction: The Reliability Contract and Full-KV Fallback
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
- Leaderboard Margins vs Hidden Model Selection: Which Gains Survive k Secret Variants?
- Greedy Decoding Is Not Precision-Invariant: BF16-vs-FP16 Flips, Margin Gating, and the Kernel-Order Axis
- LLM Prefix Caching: Where LRU Holds Up and Where It Falls Short
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- Quantization Can Change Retrieval Results While Classification Accuracy Holds
- Inside Qwen3.8-Flash-Next: Gated DeltaNet, Sparse Attention, and N-gram Memory
- Uncheatable Eval: Scoring LLMs by How Few Bits They Need
- Scale-Up vs Scale-Out: Network Bandwidth and Collective Math
hardware18 articles
- BOOST: Concurrent Host and HBM Access for LLM Decode
- CXL-SSDs For LLM Prefix Caching: Byte-Addressability Buys Nothing Without Chunk Awareness
- CXMT at 10% of DRAM Revenue: What the Figures Show
- GPU Buying Guide for LLM Inference (2026)
- Selecting the Right GPU for Qwen3 Inference
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- LLM Inference Math: From Theory to Hardware
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- Snapdragon X2 on Linux: The 80-TOPS NPU Is Not the Story — 228 GB/s Is
- AMD's MI455X 34× Throughput Claim: What the Disclosure Establishes
- AMD MI300X/MI350/MI400: More Memory, But Does the Math Win? — A Critical Silicon View of AMD Instinct vs NVIDIA
- NVIDIA NVLink: Generations, Domains and Verification
- Scale-Up vs Scale-Out: Network Bandwidth and Collective Math
deep-learning16 articles
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- Selecting the Right GPU for Qwen3 Inference
- How LLMs Actually Work — An Animated Walkthrough
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
- KV Cache Memory Growth Explorer
- LLM Inference Calculator & Break-Even Analyzer
- LLM Inference Economics: Costs & Cloud Break-Even
gpu-memory15 articles
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- Your Agent Fleet Is a Storage Array: The KV-Cache Tiering Math Behind Agentic Inference
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- HBM Capacity and KV-Cache Economics
- The KV Cache: Bit-Exact Memory Math, GQA vs MQA vs MLA, and PagedAttention
- KV-Cache Glossary: every term in the serving stack
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
- LoRA and QLoRA Fine-Tuning: The Actual Memory Math
- RAG Retrieval Math: Embedding Memory, Vector Index Size, and the Latency Budget Nobody Calculates
- UNISON and the Paused Fleet: Why Agent KV Traffic Breaks Every Cache Policy You Know
nvidia10 articles
- Qwen3-Next-80B-A3B: Architecture and Serving Limits
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
- NVIDIA Agrees to Buy Hugging Face: The Dependency Math Behind the $12.93B Headline
- NVIDIA GPU Containers on Ubuntu 24.04
- NVIDIA Error 802 on Ubuntu: Diagnose Before Trying nokaslr
- AMD MI300X/MI350/MI400: More Memory, But Does the Math Win? — A Critical Silicon View of AMD Instinct vs NVIDIA
- NVIDIA B200 vs GB200: What MLPerf Training Actually Shows
- MLPerf Inference 2026: What a Result Says About Hardware and Software
- NVIDIA NVLink: Generations, Domains and Verification
- Scale-Up vs Scale-Out: Network Bandwidth and Collective Math
kv-cache9 articles
- Random Attention vs Score-Based KV Eviction on Reasoning Traces
- BOOST: Concurrent Host and HBM Access for LLM Decode
- DeepSeek-V4.1-Flash KV-Cache Compression: What the Reported 890 Bytes per Token Includes
- FlashLoop, or: Looping a Transformer Quadruples the Bill and Then Throws Most of It Away
- HBF as the Third KV Tier: 24x Sessions or 5x Worse Latency - the Medium Is Fine, the Placement Policy Decides
- Risk-Controlled KV Eviction: The Reliability Contract and Full-KV Fallback
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- KV-Cache Tensor Decomposition: Full-Rank Head and Layer Modes in Two Models
- KVSET: LRU Stack Distance for LLM Prefix-Cache Sizing
guide8 articles
- Parakeet TDT vs Encoder–Decoder ASR: Production Trade-offs
- Selecting the Right GPU for Qwen3 Inference
- How LLMs Actually Work — An Animated Walkthrough
- Expert Parallelism: The Router's Dilemma
- Pipeline Parallelism: Pumping Data Upstream
- Tensor Parallelism: Slicing the Silicon
- NVIDIA GPU Containers on Ubuntu 24.04
- NVIDIA Error 802 on Ubuntu: Diagnose Before Trying nokaslr
datacenter7 articles
- About Florian Zimmermeister
- Parakeet TDT vs Encoder–Decoder ASR: Production Trade-offs
- Self-hosting DeepSeek V4 Pro: the actual math
- Selecting the Right GPU for Qwen3 Inference
- LLM Inference Math: From Theory to Hardware
- LLM VRAM Requirements: A Mathematical Deep Dive
- NVIDIA Groq 3 LPX: The SRAM Decode Engine and the Arithmetic NVIDIA Won't Do
vllm6 articles
- llama.cpp / Ollama vs vLLM: Which One, When, and Why
- NVIDIA Agrees to Buy Hugging Face: The Dependency Math Behind the $12.93B Headline
- The 2026 Serving-Engine Churn Tax: vLLM 0.29, SGLang 0.5.19, and TensorRT-LLM Losing TensorRT
- vLLM vs SGLang: Architecture, Overhead, and When to Pick Which
- KV Cache Memory Growth Explorer
- LLM Inference Calculator & Break-Even Analyzer
compliance5 articles
economics5 articles
- Agent-Fleet Token Economics Is Cache Economics: Decomposing the 60x Dollar Spread
- The Cache-Read Price War of September 2026 That Wasn't Three-Sided
- Who Pays for the KV Cache? Metering Rules Are a Pricing Decision in Disguise
- KITE: Scale the Model, Freeze the KV Bill — the Prefill Invariant arXiv:2609.27294 Actually Proves (and the Three Things It Does Not)
- LLM Inference Economics: Costs & Cloud Break-Even
moe5 articles
ai-act4 articles
optimization4 articles
quantization4 articles
- Disaggregated Quantization: Splitting the Prefill and Decode Price System Across Two Formats
- PTQ Configuration Pricing: Predicting the Quantization Price Before You Build the Model
- LLM Quantization: Bit Layouts, Block Scales, and the VRAM Math
- Quantization Can Change Retrieval Results While Classification Accuracy Holds