Unlocking Efficiency in AI: Fireworks Ember-1 Cuts Token Usage by 40%


Unlocking Efficiency in AI: Fireworks Ember-1 Cuts Token Usage by 40%

Fireworks AI has released Ember-1, a post-trained version of Moonshot AI’s Kimi K3 that uses about 40% fewer tokens while trying to keep the same level of quality. Instead of simply telling the model to “think less” at runtime, Fireworks trained it to reason in a shorter and smarter way, which matters a lot for coding, tool use, and long agent tasks. Ember-1 is available as an API-only Research Preview through Fireworks Serverless, and while benchmark tests and customer A/B results look strong, the model weights, training code, and exact training methods have not been released.

Why Ember-1 Matters for Home Users and Developers

  • Ember-1 matters because many reasoning models are like a student who writes three full notebook pages before giving a short answer to one math question.
  • That long thinking process may look smart, but it also costs money because every extra token has to be generated, stored, and sometimes read again in the next step.
  • Fireworks AI says Kimi K3 can spend more than 90% of its generated tokens on internal reasoning, which becomes a real problem in long multi-turn agent workflows.
  • Think of it like asking a friend for directions, but before telling you “turn left,” they repeat the whole history of every road they have ever taken.
  • In simple chat tasks, that may only feel annoying, but in coding agents, search agents, and tool-using systems, it becomes expensive very fast.
  • If an AI agent takes 20 or 25 steps to finish one software task, it may keep dragging its old thoughts into every new step.
  • That means early reasoning gets replayed again and again, which increases context size and cost.
  • Fireworks built Ember-1 to fix this exact problem without badly hurting task performance.
  • This is an important detail because lowering the reasoning setting is not the same thing as training a model to reason better.
  • Lower settings often act like telling a runner to move slower and hope they still win the race.
  • Post-training for efficiency is more like teaching the runner a better path so they use less energy and still reach the finish line at nearly the same speed.
  • For developers building agents, this can mean fewer wasted tokens, smaller bills, and smoother production use.
  • For teams watching every API dollar, a model that keeps quality while speaking less can feel like replacing a car that burns too much fuel with one that goes just as far on less gas.
  • That is why Ember-1 is not just another model launch but a direct answer to the growing “token tax” in practical AI systems.

How Fireworks Used Research to Train Shorter Reasoning

  • Fireworks says Ember-1 was created by post-training Moonshot AI’s open-weight Kimi K3, which means the base model already existed and then got extra training for a more specific goal.
  • That goal was not to remove thinking completely, because not all reasoning is useless.
  • Some reflection helps a model notice mistakes, react to tool feedback, and update its plan during a hard task.
  • So the research challenge was to cut wasted loops while keeping useful self-checking.
  • Imagine cleaning a backpack before school.
  • You do not throw away your books, but you do remove broken pens, snack wrappers, and random papers you do not need.
  • Ember-1 seems to follow that same idea by keeping helpful reasoning and dropping repeated or unproductive chains of thought.
  • According to Fireworks, the training collection covered math, coding, instruction following, conversation, search, tool use, and software engineering.
  • That matters because shorter reasoning in only one narrow area would not be enough for real-world use.
  • A coding agent may need to read docs, call tools, check outputs, and then explain results to a user in plain language.
  • Fireworks also said the work included both standalone tasks and longer multi-step interactions.
  • That is important because long tasks are where token costs grow the fastest.
  • The company ran more than 50 training experiments and over 200 evaluations, which suggests the final result came from repeated testing and adjustment rather than a single lucky run.
  • Task and environment feedback were used for on-policy planning and learning, meaning the model learned while reacting to what happened during tasks.
  • This is a bit like learning to ride a bike by making small balance fixes while moving, instead of only reading a manual about bicycles.
  • Fireworks also said it created new training algorithms, but those details have not been published.
  • All training ran on Fireworks Serverless Training, and the company says it used its own data rather than customer data.
  • That privacy point may matter for enterprise users who worry about their internal code or business data being reused in training.
  • Still, because the weights and code are not public, outside researchers cannot yet reproduce the method or inspect it deeply.
  • So the training story is promising, but some parts remain closed behind the company’s research wall.

What the Benchmarks Say Across Tutorials, Coding, and Tool Tasks

  • The benchmark results are one of the strongest reasons people are paying attention to Ember-1.
  • Fireworks compared Ember-1 against Kimi K3 at low, high, and max reasoning effort levels.
  • On Terminal Bench 2.1, Ember-1 scored 82.0%, which beat K3 Max at 80.9%.
  • On DeepSWE 1.1, Ember-1 scored 75.2%, which was also better than K3 Max at 66.4%.
  • These are meaningful wins because they show Ember-1 was not only cheaper in token use but also stronger on some difficult engineering-style tasks.
  • At the same time, the model was not best on every benchmark.
  • On SWE-bench Verified, Ember-1 scored 92.2%, slightly below K3 Max at 93.2%.
  • On SWE-Interact, Ember-1 reached 20.0%, also slightly below K3 Max at 21.3%.
  • This matters because it shows the results are more believable than a claim that says “we win everywhere.”
  • Real AI systems often improve one thing while giving up a little somewhere else.
  • What makes Ember-1 interesting is that the tradeoff appears small while the token savings are large.
  • Fireworks says reasoning was shortened by 35% to 50% across seven benchmarks and two customers’ production traffic without sacrificing accuracy.
  • That is a big statement because in AI, cheaper usually comes with weaker quality.
  • If those numbers hold in wider use, Ember-1 could become a useful option for teams that need a model to solve complex tasks many times per day.
  • Fireworks also highlighted Doximity’s Bedside Bench, a physician-validated set of 500 clinical cases, where Ember-1 reportedly reached a new cost-per-task Pareto frontier.
  • In simple words, that means it gave a very strong balance between quality and cost.
  • For SEO readers searching terms like “Ember-1 benchmark results,” “Kimi K3 vs Ember-1,” or “reasoning model token efficiency,” these numbers are the heart of the story.
  • They suggest Ember-1 is not merely shorter in output but more disciplined in how it spends its reasoning budget.
  • It is like seeing two students take the same test, but one writes a cleaner answer sheet, finishes sooner, and still gets almost the same or even a better score.
  • That kind of efficiency is exactly what many AI product teams have been waiting for.

API Access, Pricing, and the Real Cost Story Behind Voice AI Era Workloads

  • One of the most practical parts of this release is the pricing story.
  • Ember-1 costs the same per token as Kimi K3 on Fireworks: $3.00 per 1M input tokens, $0.30 per 1M cached input tokens, and $15.00 per 1M output tokens.
  • At first that may sound like there is no price benefit at all, but the savings come from using fewer tokens in the first place.
  • That is a bit like two printers charging the same price per page, but one printer needs only 6 pages to explain something while the other needs 10 pages.
  • In the published A/B example, output tokens dropped from 49.3K to 29.9K per task.
  • Reasoning tokens fell by 71.3%, and total tokens fell by 39%.
  • The task score was almost unchanged at 0.753 for Ember-1 versus 0.751 for K3.
  • Average steps also decreased from 23.8 to 21.4.
  • Using Fireworks pricing, the output-only cost works out to about $0.74 versus $0.45 per task.
  • That may not sound huge for one task, but in production it adds up quickly.
  • If a company runs 100,000 tasks, that difference becomes tens of thousands of dollars.
  • This is why token efficiency is becoming a business issue, not just a model design issue.
  • Fireworks says Ember-1 is already being used in production by at least one customer.
  • However, deployment options are limited because Ember-1 is only available as an API-only Research Preview through Fireworks Serverless.
  • It is also available through OpenRouter using the Fireworks endpoint at the same pricing.
  • There is no self-hosting option right now because the model weights, training code, and exact learning methods have not been released.
  • That means companies that need full control, offline use, or deep model inspection may hesitate.
  • For those teams, the lack of open release is like being allowed to drive a fast car only on the maker’s track, not in your own garage.
  • But for teams that already prefer managed APIs, Ember-1 may still be a very attractive option, especially for coding, support agents, and complex workflows where each saved token lowers the total bill.

Robotics Lessons from Ember-1 for the Future of Agentic AI

  • Even though Ember-1 is not a robotics model, its biggest lesson fits robotics and agentic AI very well: smart systems need to act with focus, not just with more words.
  • In robotics, a machine that pauses too long, rechecks everything too many times, or repeats old steps can waste power and time.
  • The same thing happens with software agents that overthink before taking the next action.
  • A useful agent should be like a good mechanic who checks the engine carefully, but does not take the whole car apart every time the oil light turns on.
  • Ember-1 shows that AI progress is no longer only about getting higher scores.
  • It is also about using compute more wisely, especially for long chains of actions.
  • This matters for agentic AI, where the model may search the web, open tools, edit code, run tests, read logs, and answer the user all in one session.
  • Each extra reasoning block becomes part of the next prompt, which makes future calls heavier.
  • That snowball effect is why Fireworks focused on shorter reasoning traces instead of just fast answers.
  • A shorter trace can make the whole session lighter from start to finish.
  • Picture a hiker carrying supplies up a hill.
  • If the backpack gets bigger at every stop, the trip becomes harder and slower.
  • If the hiker keeps only what is useful, the full journey becomes easier.
  • That is very close to what Ember-1 tries to do inside long AI tasks.
  • For the future, this could push more labs to train models for efficient reasoning instead of only maximizing raw thinking length.
  • It may also influence how AI infrastructure companies measure value, because users increasingly care about cost per finished task, not just benchmark glory.
  • Still, people should remember that Fireworks published these results itself, so wider outside testing will be important.
  • Until more independent evaluations appear, Ember-1 should be seen as a very promising research preview rather than a final answer for every workflow.
  • Even so, the model points toward a simple but powerful idea: in AI, better thinking may sometimes mean thinking with less waste.

Conclusion

Ember-1 shows a clear new direction for reasoning models: keep the brains, cut the waste. Fireworks AI says the model delivers Kimi K3-like quality with about 40% fewer tokens by using post-training instead of just lowering reasoning effort at runtime. Benchmark scores, production A/B tests, and cost-per-task results suggest this approach may help developers run coding and agentic workloads more efficiently. The biggest limitation is access, because Ember-1 is still an API-only Research Preview with no released weights or training code. Even so, Ember-1 makes one thing very clear: in the next stage of AI, efficiency may matter just as much as raw intelligence.

Source: https://www.marktechpost.com/2026/09/28/fireworks-ai-releases-ember-1-a-post-trained-kimi-k3-that-uses-about-40-fewer-tokens/

Post a Comment

Previous Post Next Post