GPU Kernel Optimization
Parent: ai-inference-engines
Related: inferact, radixark
Overview
GPU kernel optimization is the discipline of improving the efficiency of low-level compute kernels that run on GPU hardware during model inference. It's the foundational layer beneath inference engines like vLLM and SGLang.
How GPUs Run Models
When model.generate() is called, the GPU executes hundreds to thousands of kernels — small parallel functions such as:
- GEMM (General Matrix Multiplication) — the heavy lifter
- FlashAttention — efficient attention computation
- LayerNorm — normalization
- Activation functions (ReLU, GELU, SiGLU)
Bottleneck insight: GPUs are rarely compute-bound — they're usually memory bandwidth-bound. The core optimization challenge is minimizing data movement between HBM (High Bandwidth Memory) and compute units.
Key Optimization Techniques
1. Memory Access Optimization
- Kernel fusion: Combine multiple operations into a single kernel to reduce HBM reads/writes
- Triton compiler: Write fused kernels in Python-like syntax, compile to optimized CUDA
- FlashAttention: IO-aware attention that reduces HBM traffic by ~10x
2. Quantization Kernels
| Format | Precision | Speedup vs FP16 | Quality Impact |
|---|---|---|---|
| INT8 | 8-bit | ~2x | Minimal |
| INT4 | 4-bit | ~4x | Moderate |
| FP8 (Hopper) | 8-bit float | ~2-3x | Very low |
3. Tensor Parallelism Kernels
- Split model weights across multiple GPUs
- NVLink/NVSwitch for high-bandwidth inter-GPU communication
- Critical for 70B+ models that don't fit on single GPU
4. FlashDecoding++ (无问芯穹)
Novel kernel-level optimization for the decode phase:
- vs Hugging Face: 4.86x speedup (A100)
- vs FlashDecoding: 1.37x speedup
- vs vLLM: 1.24x speedup in decode stage
Inference Engine Comparison (Kernel-Level)
| Engine | Quantization | Fusion | Multi-GPU | Custom Kernels |
|---|---|---|---|---|
| vLLM | AWQ, GPTQ, GGML | Yes | TP, PP | PagedAttention |
| SGLang | AWQ, GPTQ | Yes | TP, PP, EP | RadixAttention |
| LMDeploy | INT4, INT8 (TurboMind) | Yes (C++) | TP | TurboMind engine |
| TensorRT-LLM | FP8, INT8 | Yes | TP, PP | Highly optimized |
| Inferact | (uses vLLM) | Yes | Yes | vLLM kernel stack |
| RadixArk | (uses SGLang) | Yes | Yes | SGLang kernel stack |
Why It Matters
- H100 cluster cost: ~$3-5/hr per GPU
- Kernel optimization directly = lower inference cost per token
- For 1B+ token/day deployments: 2x kernel speedup = $millions saved annually
Novel Attention Kernels: Open vs Closed Backprop
Chinese frontier labs have introduced custom attention mechanisms that reduce KV cache memory pressure at long context lengths. Their GPU kernel openness varies, with implications for downstream training.
| Mechanism | Lab | Inference Kernel | Training (Backprop) Kernel | Status |
|---|---|---|---|---|
| MLA (Multi-head Latent Attention) | deepseek | Open (in vLLM/SGLang) | Partially open | Inference widely deployed; full training kernels not in public repos |
| KDA (Key-Domain Attention) | kimi (K3) | Open (FlashKDA) | Closed / not released | Only the forward pass is open-source; backprop kernel is missing, preventing full training reproducibility |
| FlashAttention | Dao et al. | Open (BSD) | Open (BSD) | Reference implementation — fully open |
Implication for downstream users: Standard LoRA/QLoRA/SFT fine-tuning on these models works via PyTorch + Hugging Face Transformers (which fall back to standard attention during training), but reproducing or modifying the original pre-training pipeline — including the forward+backward pass at native speed — requires reverse-engineering the training kernels. This is a practical barrier to full training-stack reproducibility, even when model weights are open [local: 2026-08-03-ai-models.md].
Sources
- vLLM paper (PagedAttention)
- FlashAttention paper (Dao et al.)
- 无问芯穹 FlashDecoding++ technical disclosures
- local: 2026-08-03-ai-models.md