Engineering blog

GPU inference, written down

Technical notes on scheduling algorithms, KV cache behavior, quantization interactions, and the day-to-day reality of running open-weight models on heterogeneous GPU fleets.

Run inference on your own fleet

If the scheduling problems described here apply to your infrastructure, we work directly with early access partners on fleet-specific configuration.