# Ember-1: How Fireworks AI Is Teaching Models to Reason More Efficiently
## A New Approach to Cost-Efficient Reasoning
A new model called Ember-1 has entered the AI landscape as a specialized variant of Kimi K3, developed by Fireworks Research. Unlike conventional approaches that simply dial down reasoning effort at inference time, Ember-1 was built from the ground up to be a more efficient thinker. It achieves this through a post-training process that preserves task quality while dramatically cutting the number of tokens consumed during reasoning.
The core innovation is straightforward in principle but demanding in execution: train a model to identify which parts of its own reasoning are valuable and which are redundant, then learn to produce shorter reasoning traces without sacrificing accuracy.
## Why Shortening Reasoning Traces Matters
Large reasoning models have become increasingly capable, but they come with a hidden cost — they tend to generate far more tokens than necessary. Industry observations suggest that some reasoning models spend more than nine-tenths of their output tokens on internal deliberation rather than delivering final answers.
This problem escalates quickly in real-world scenarios. When models are used in multi-turn agentic workflows, each successive turn requires the model to re-read the reasoning traces from earlier turns. Because context length grows roughly in proportion to the square of the number of turns, early verbose reasoning gets replayed and re-billed on every future interaction. For teams running production coding assistants or multi-step analysis pipelines, this creates a compounding cost burden that no amount of inference-time parameter tweaking can fully address.
Fireworks AI found that simply reducing the reasoning effort slider on Kimi K3 did not resolve the issue — lower effort settings degraded quality too significantly to be practical. The solution was to retrain the model itself to reason more efficiently.
## How Ember-1 Was Built
The development process started with Kimi K3, an open-weight model from Moonshot AI. Through extensive post-training, the research team shaped Ember-1 into a model that keeps the useful habits of deep reasoning while shedding wasteful patterns.
Not all thinking is equal. Ember-1 was trained to preserve the kind of self-reflection that actually helps — such as revisiting an assumption mid-solve or incorporating user feedback to correct course. At the same time, it learned to cut redundant restatement and unproductive loops that don’t contribute to the final output.
The training data spanned a wide range of domains including mathematics, software engineering, coding, instruction following, conversation, web search, and tool use. Both standalone problems and extended multi-step interactions were included. Task outcomes and environment feedback guided the model’s on-policy learning throughout the process.
The team ran over fifty distinct training experiments and conducted more than two hundred evaluations. They also developed novel training algorithms that have not been publicly released. All training was performed using Fireworks’ own serverless infrastructure, and the company confirmed that no customer data was used in the process.
## Performance: Shorter Traces, Equal or Better Accuracy
Benchmark results paint a compelling picture. Fireworks compared Ember-1 against Kimi K3 at three different reasoning effort levels across multiple standard evaluations. Cost calculations were based on publicly available Kimi K3 API pricing.
Across the board, Ember-1 delivered results that matched or exceeded the highest-effort version of Kimi K3 while generating substantially fewer tokens. On Terminal Bench 2.1, Ember-1 achieved 82.0% accuracy compared to Kimi K3 Max’s 80.9%, at a reported 51.9% lower cost. On DeepSWE 1.1, Ember-1 reached 75.2% versus K3 Max’s 66.4%, with a 23.7% reduction in cost.
SWE-bench Verified and SWE-Interact results were slightly closer between the models, but Ember-1 still came in competitive. On the τ-2 Bench Airline benchmark, it edged out all K3 configurations with 66% accuracy.
## Real-World Production Results
Beyond benchmarks, Fireworks ran live A/B tests with two customers on production coding workloads. The results mirrored the controlled evaluations closely. Output tokens per task dropped from approximately 49,300 to 29,900 — a reduction of roughly 39%. Reasoning tokens specifically fell by 71.3%. Average task steps decreased from 23.8 to 21.4.
Critically, task quality scores remained essentially unchanged: Ember-1 scored 0.753 compared to Kimi K3’s 0.751. One of the two customers has already moved Ember-1 into full production use.
## Pricing and Deployment
Ember-1 is priced identically to Kimi K3 on the Fireworks platform: $3.00 per million input tokens, $0.30 per million cached input tokens, and $15.00 per million output tokens. The savings are entirely driven by generating fewer tokens rather than a discounted rate.
Deployment is available through the Fireworks serverless API in a Research Preview capacity. The model weights, training code, and detailed training algorithms have not been released publicly, meaning self-hosting is not currently possible.
## Frequently Asked Questions
**What is Ember-1?**
Ember-1 is a specialized reasoning model developed by Fireworks Research through post-training of Kimi K3. It is designed to produce shorter reasoning traces while maintaining or improving task accuracy compared to the base model.
**How is it different from simply lowering the reasoning effort setting?**
Reducing reasoning effort at inference time trades quality for efficiency. Ember-1 avoids this trade-off by being fundamentally retrained to reason more efficiently — keeping what is useful and eliminating what is redundant.
**Can I self-host Ember-1?**
Not currently. Ember-1 is only available through the Fireworks serverless API as a Research Preview. The weights and training code have not been released.
**What kind of cost savings does it offer?**
In production A/B testing, token usage dropped by approximately 39%, with reasoning tokens specifically reduced by over 70%. Since the per-token price is the same as Kimi K3, the savings are entirely from using fewer tokens.
**Does it sacrifice quality?**
On most benchmarks, Ember-1 matches or exceeds the highest-effort version of Kimi K3. In live production tests, quality scores were effectively unchanged (0.753 vs 0.751).
**What domains was it trained on?**
Mathematics, coding, software engineering, instruction following, conversation, web search, and tool use — covering both single standalone tasks and extended multi-step interactions.
## Conclusion
Ember-1 represents an important shift in how we think about reasoning in large language models. Instead of accepting verbose thinking as an unavoidable byproduct of intelligence, Fireworks AI has demonstrated that a model can be trained to trim the fat while keeping the substance. For teams deploying agentic workflows and multi-turn applications, the implications are significant — lower costs, faster responses, and no compromise on quality. As the industry moves toward more production-heavy AI workloads, efficiency-focused training approaches like this are likely to become a central theme in model development.
Thank you for reading



