Early Access Pricing
Node-based licensing. No per-token tax.
Inferact charges by GPU node, not by token volume. Your traffic spikes are your business; the scheduler's job is to handle them without charging you extra for doing its job well.
Up to 8 GPU nodes
- Continuous batching engine
- Prefix-aware KV cache
- Prometheus metrics endpoint
- Community Slack support
- Quickstart documentation
Up to 32 GPU nodes
- Everything in Hobbyist
- Quantization-aware dispatch
- Speculative decoding integration
- OOM recovery and pressure guard
- Priority queue with SLA tiers
- Email + Slack dedicated support
Unlimited nodes
- Everything in Team
- Multi-node tensor parallelism
- Custom SLA configuration
- On-premise deployment support
- Dedicated engineering contact
- Custom integration work
What is included
Across all tiers
| Capability | Hobbyist | Team | Fleet |
|---|---|---|---|
| Continuous batching | |||
| Prefix-aware KV cache | |||
| Prometheus metrics | |||
| Quantization-aware dispatch | - | ||
| Speculative decoding | - | ||
| OOM recovery | - | ||
| Multi-node tensor parallelism | - | - | |
| Dedicated engineering support | - | - |
Common questions
Pricing FAQ
Each GPU accelerator in your fleet counts as one node. A 4x H100 server counts as four nodes. You pay a flat monthly fee per node regardless of how many tokens pass through it. Early access pricing is not public yet; contact us for specifics based on your fleet size.
Early access participants work with us directly to run an initial profiling session on their cluster. We use that session to characterize your GPU mix and show projected throughput before you commit. Reach out through the contact form with your hardware configuration.
Yes. Inferact works on bare-metal and on cloud instances (AWS p4d/p5, Azure NDv4/NDv5, GCP A3). The runtime runs alongside your existing workloads and does not require special host privileges beyond CUDA device access.
You are licensed per node. Adding GPUs to your cluster means updating your node count at the next billing cycle. There is no long-term contract required during early access; terms are month-to-month.
Yes. Inferact schedules across all models served from a node. If a single H100 serves both Llama-3 70B shard traffic and a smaller 8B model, that is still one node in your license count. The scheduler handles cross-model routing and KV cache partitioning automatically.
CUDA 11.8 or later, Python 3.10 or later, and Linux (Ubuntu 20.04+ or similar). The scheduler daemon requires approximately 2GB of system RAM per 8 GPU nodes. Kubernetes support via Helm chart is documented in the quickstart.
Have a fleet that does not fit the tiers above?
Unusual GPU mixes, air-gapped deployments, and very large multi-site fleets are situations we want to hear about. Contact engineering directly.