# Prime Inference: A New Serving Platform for Open-Source Frontier Models
The landscape of open-source AI infrastructure is shifting rapidly as new platforms emerge to bridge the gap between experimental models and production workloads. One such development comes from Prime Intellect, which has introduced Prime Inference — a dedicated serving platform designed to handle the demands of frontier open-source models at scale. With serverless endpoints, reserved capacity options, and multi-datacenter deployment, the platform aims to make high-performance model serving accessible to a broader developer community.
## Understanding the Platform
Prime Inference functions as the serving layer within Prime Intellect’s broader open training ecosystem. The company has already established itself in the post-training space with tools like prime-rl, verifiers, and sandboxes. By adding a robust serving layer, the company completes the development-to-deployment pipeline, creating a feedback loop where production traces from deployed models can flow back into training processes.
Before its public release, the platform processed an extraordinary volume of internal traffic — nearly a trillion tokens per day. This traffic was generated by reinforcement learning rollouts, synthetic data generation pipelines, rigorous evaluations, and long-running coding agents. The scale of this internal testing suggests the platform has been stress-tested under conditions that mirror real-world production demands.
According to the company, its GLM-5.3 endpoint ranks among the fastest available on major open-source serving rankings. The platform also reports a near-zero tool-call error rate and has maintained 100% uptime since its launch — metrics that speak to the reliability of the underlying infrastructure.
## Key Features and Capabilities
Prime Inference offers two distinct operating modes to accommodate different workload patterns. Serverless endpoints handle variable, unpredictable demand by automatically scaling resources up and down based on incoming traffic. Reserved capacity, on the other hand, provides dedicated GPU resources for sustained workloads that require consistent performance and predictable costs.
The platform is designed to be familiar to developers already working with major AI providers. It is fully OpenAI compatible, meaning developers can point any existing OpenAI SDK directly at the Prime Inference endpoint without modifying their application code. This compatibility lowers the barrier to adoption significantly.
Reliability is built into the architecture through automatic failover mechanisms. Traffic is routed across multiple datacenters, and if one deployment experiences issues, requests are automatically redirected to healthy instances. The underlying hardware currently utilizes NVIDIA Blackwell GPUs, with NVIDIA’s next-generation Vera Rubin architecture listed as an upcoming upgrade.
Billing is handled through a unified system that includes team-level usage tracking. While per-model pricing details have not yet been fully published in the documentation, the billing infrastructure is designed to provide transparency and control for teams managing multiple models.
## The Technical Architecture
The serving stack represents a carefully curated combination of several leading open-source technologies. It integrates NVIDIA Dynamo for efficient routing and load balancing, vLLM for high-throughput model inference, Mooncake for additional caching layers, and FlashInfer for optimized attention computation. The platform was developed in collaboration with Inferact and NVIDIA, with improvements contributed back to the upstream projects.
The target workload profile is distinctly agentic in nature. A typical agent interaction adds approximately 6,000 tokens to a base prompt that already contains around 140,000 tokens of context. The platform benchmarks this workload pattern using SemiAnalysis AgentX metrics and incorporates cold arrival scenarios to simulate real-world traffic conditions.
### Prefill and Decode Disaggregation
One of the platform’s most significant architectural decisions is the separation of prefill and decode operations onto different GPU groups. NVIDIA Dynamo handles routing between these groups, while vLLM runs the model on each one. During the decode phase, the decoder nodes pull precomputed key-value (KV) data through NIXL, NVIDIA’s high-performance interconnect library.
In testing, this disaggregated approach yielded nearly a 40% reduction in p90 inter-token latency compared to configurations where prefill and decode share the same GPU resources. This improvement is particularly meaningful for interactive applications where user experience depends on low-latency token generation.
### Cache-Aware Routing
Dynamo’s KV-aware router makes intelligent decisions by weighing the degree of cached prefix overlap against the current queue depth on each worker. For returning sessions, this means the system can identify and reuse the decoder node that already holds the relevant KV cache, minimizing redundant computation. Mooncake layers on top of this by providing a second KV cache tier stored in host DRAM, expanding the effective cache capacity beyond what GPU memory alone can offer.
## Performance Benchmarks: GLM-5.3 on GB200 NVL72
The platform’s performance characteristics were measured using the GLM-5.3 model running on NVIDIA GB200 NVL72 hardware, with an interactivity target of 100 end-to-end tokens per second per user. The optimal operating point was found at a 1:4 ratio of prefill to decode resources.
At this configuration, the system achieved 66 concurrent sessions per prefill group while maintaining 101 tokens per second per user and 100 output tokens per second per GPU. Several architectural optimizations contributed to these results.
The DEP8 prefill topology demonstrated roughly five times more usable prefix-cache capacity compared to the TEP8 configuration on identical hardware. By reducing the prefill token budget from 8,000 to 4,000 tokens per GPU per processing step, median queue wait times dropped dramatically from 550 milliseconds to just 110 milliseconds. Median time to first token also improved by approximately 20%.
NVFP4 KV compression proved especially impactful for memory efficiency. Each multi-head latent attention (MLA) cache row was reduced from 576 bytes to 352 bytes, allowing cached tokens per decoder to increase from 1.09 million to 1.63 million — a roughly 50% improvement in cache capacity within the same memory footprint. A native sparse-MLA kernel built into the system achieves approximately 12.0 microseconds at 15 query tokens, outperforming both staged implementations at 17.7 microseconds and FP8 variants at 13.7 microseconds.
The BLHNC (block-major, 64-token blocks) KV layout reduced transfer descriptors from 19,559 to approximately 1,940, bringing mean transfer time down from 146 milliseconds to 78 milliseconds. This optimization keeps a block’s data contiguous across many layers, enabling NIXL to issue fewer but larger, more efficient data transfers.
## Tool Call Reliability Improvements
Agent reliability depends heavily on accurate tool call formatting. Prime Intellect’s engineering team contributed a structural-tag builder to Dynamo that properly formats GLM’s tool call schema. vLLM then uses xgrammar to mask tokens that would violate the expected tool format, preventing malformed requests from reaching the model. The team also resolved several parsing bugs, including one where the less-than character was being incorrectly decoded inside code blocks.
These improvements contribute to the platform’s reported near-zero tool-call error rate, which is critical for agentic workloads where incorrect tool calls can cascade into broader failures.
## FAQ
**What types of models can be served on Prime Inference?**
Prime Inference is designed to serve frontier open-source models, with GLM-5.3 as a currently featured model. The platform’s architecture is general-purpose enough to support other open-source models, though specific model support may evolve over time.
**How does serverless mode differ from reserved capacity?**
Serverless mode automatically scales compute resources based on incoming request volume, making it ideal for workloads with variable or unpredictable traffic patterns. Reserved capacity dedicates specific GPU resources to your workload, providing consistent performance and more predictable costs for sustained, high-volume operations.
**What hardware is currently used, and what is coming next?**
The platform currently runs on NVIDIA Blackwell GPUs deployed across GB200 NVL72 systems. NVIDIA’s Vera Rubin architecture has been announced as the next-generation hardware upgrade.
**Is the platform OpenAI API compatible?**
Yes. Prime Inference is fully OpenAI compatible, meaning any application using the OpenAI SDK can be pointed to Prime Inference endpoints with minimal or no code changes.
**What is the expected uptime and reliability?**
The platform has achieved 100% uptime since launch and implements automatic failover across datacenters, routing traffic away from unhealthy deployments to maintain availability.
**How is billing structured?**
Prime Inference uses a unified billing system with team-level usage tracking. Per-model pricing details are still being finalized and are not yet fully published in the public documentation.
**What was the internal testing scale before public launch?**
Before opening to public users, the platform processed nearly one trillion tokens per day internally, generated through RL rollouts, synthetic data generation, evaluations, and long-running coding agents.
**Can deployed models feed back into training?**
Yes. Prime Inference is designed as the serving layer of an open training stack. Production traces generated by deployed models can be captured and used to improve future training iterations, closing the loop between deployment and model improvement.
## Conclusion
Prime Inference represents a significant step forward in the infrastructure supporting open-source frontier models. By combining serverless flexibility with reserved capacity options, OpenAI-compatible APIs, and a carefully engineered technical stack spanning Dynamo, vLLM, Mooncake, and FlashInfer, the platform addresses many of the operational challenges that come with deploying large language models in production environments.
The performance benchmarks, particularly the disaggregated prefill/decode architecture and the innovative NVFP4 KV compression techniques, demonstrate that substantial improvements in latency and throughput are achievable with thoughtful systems engineering. The commitment to contributing improvements back to upstream projects also reflects a collaborative approach that benefits the broader open-source AI community.
As agentic workloads continue to grow in complexity and prevalence, platforms like Prime Inference that are specifically architected for these demands will play an increasingly important role in the AI ecosystem. The combination of reliability, performance optimization, and developer-friendly design positions the platform as a compelling option for teams looking to deploy frontier open-source models at scale.
Thank you for reading



