Engineering blog
GPU inference, written down
Technical notes on scheduling algorithms, KV cache behavior, quantization interactions, and the day-to-day reality of running open-weight models on heterogeneous GPU fleets.
Scheduling
Memory
KV Cache Eviction Policies Compared: LRU, Prefix-Aware, and Adaptive
Hardware
Running Llama-3 8B on RTX 4090: PagedAttention Configuration Notes
Batching
Continuous Batching vs Static Batching: The Throughput-Latency Tradeoff
Quantization
Quantization-Aware Scheduling: Mixing INT4 and FP16 in the Same Fleet
Multi-Node
Multi-Node Tensor Parallelism: When NVLink Isn't Enough
Memory
OOM Recovery and Memory Pressure in Long-Context Inference
Batching
Speculative Decoding and Batching: Why They Conflict and How to Manage It
Benchmarks
vLLM, TGI, and Inferact: Throughput Benchmark on Llama-3 70B
Performance
CUDA Graph Replay: Eliminating Kernel Launch Overhead for Inference
Memory
Flash Attention v3 and KV Cache Bandwidth: What Changes for Your Setup
Scheduling
A100 and H100 in the Same Fleet: Scheduling Without Waste
Architecture
Prefill-Decode Disaggregation: Separating Compute Phases for Better GPU Utilization
Run inference on your own fleet
If the scheduling problems described here apply to your infrastructure, we work directly with early access partners on fleet-specific configuration.