Fireworks Ember-1 Review (9/9, $4.63): Cheaper Per Token Than Opus 5.5, Dearer Per Task
Fireworks Ember-1 is a model whose whole pitch is spending fewer tokens, and on our executed Python benchmark it did the job: 9 out of 9 at $4.63 per 1,000 tasks (priced 2026-10-02) with a 4.1-second mean. The surprise is the comparison it loses. Ember-1 lists at $3 in / $15 out per 1M, a quarter below Claude Opus 5.5 at $4 / $20. Measured the same day on the same nine prompts, Opus 5.5 cost $4.03 and Ember-1 cost $4.63, both priced 2026-10-02 — because Ember-1 wrote 54% more output tokens. A cheaper rate card did not survive contact with the token count.
Fireworks is best known as the place you run other labs' open-weight models. Ember-1 ships under its own name, as fireworks/ember-1 — though, as Fireworks itself says, it is another lab's model underneath. That makes it a useful test of a narrow question: when an inference host tunes a model for its own economics, does the saving reach your bill?
The result
| Metric | Fireworks Ember-1 |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $4.63 (priced 2026-10-02) |
| Mean latency | 4.1s |
| Reasoning tokens per call | 166 |
| Tokens across the suite | 1,341 in / 2,512 out |
| List price in / out | $3 / $15 per 1M (2026-10-02) |
| Context window | 1,048,576 |
| Measured | 2026-10-02 |
Among the 75 models in our set that score 9/9, Ember-1 ranks 53rd cheapest and 14th fastest. That is the profile of a quick, mid-premium model: near the fast end, past the midpoint on cost.
What Ember-1 is, according to Fireworks
Everything in this section is third-party — Fireworks' own claims and listings, read 2026-10-02 — not something we measured.
- Origin. Fireworks' announcement, dated 2026-09-23, describes Ember-1 as a specialized model from Fireworks Research built on Moonshot AI's Kimi K3. It is a post-train of an existing base, not a new architecture.
- The claim. The same post says Ember-1 delivers Kimi K3's quality with roughly 40% fewer tokens, by training out redundant reasoning loops. The evidence is Fireworks' own benchmark runs and A/B tests; independent coverage has noted that it is vendor-run.
- Status. Fireworks frames it as a research model with time-limited serverless access, and Vercel's AI Gateway listing calls it a research preview. Treat availability as provisional.
- Price. Vercel lists $3 in / $15 out per 1M, matching what we priced on 2026-10-02 — and Fireworks' post gives the same rate card for Kimi K3 on its platform, so any saving has to come from token count, not from the rate.
Two things we deliberately leave out. Several aggregator pages repeat a parameter count, but Fireworks' announcement does not state one, so neither do we. And we found no licence statement for the Ember-1 weights themselves; that Kimi K3 is open-weight does not tell you Ember-1 is. Our own launch-day look at the base model is in the Kimi K3 review.
Cheaper per token, dearer per task
Anthropic's two models new to this sweep went through the same nine prompts on the same day, which gives an unusually clean comparison:
| Model | List in / out per 1M | Output tokens (suite) | Reasoning / call | Cost / 1k tasks | Latency |
|---|---|---|---|---|---|
| Fireworks Ember-1 | $3 / $15 | 2,512 | 166 | $4.63 | 4.1s |
| Claude Opus 5.5 | $4 / $20 | 1,630 | 33 | $4.03 | 4.8s |
| Claude Sonnet 5.5 | $2 / $10 | 1,569 | 31 | $1.95 | 3.4s |
All three scored 9/9; all prices and costs are as of 2026-10-02. Ember-1's rate card is exactly 0.75 of Opus 5.5's on both input and output ($3 / $4, $15 / $20). Yet the measured bill runs the other way: $4.63 against $4.03, 60 cents more per thousand tasks, both priced 2026-10-02. The reason is in the token columns. Ember-1 emitted 2,512 output tokens across the suite to Opus 5.5's 1,630 — 1.54 times as many — and about five times the reasoning tokens per call, 166 against 33. At $15 per million, output is the expensive side of Ember-1's card, five times its input rate, so extra output tokens land directly on the bill.
Ember-1 does buy something for that 60 cents: it was 0.7 seconds faster than Opus 5.5 at the mean. Against Claude Sonnet 5.5, though, it buys nothing on this suite. Sonnet 5.5 scored the same 9/9 at $1.95 (priced 2026-10-02), 0.7 seconds faster than Ember-1, so Ember-1 costs 2.4 times as much for a slower answer. This is the lesson of our piece on reasoning tokens deciding the bill again: the rate card is an input, the token count is the multiplier, and only the product matters.
The $4 to $5 band
To place Ember-1 among its price peers we took the eight 9/9 models in our fact sheet with a full entry and a measured cost between $4.03 and $4.99 per 1,000 tasks (capture dates in the table). This is a local set chosen by cost, not every model at that price, and each cost carries the date its list price was captured:
| Model | Cost / 1k | Priced on | Latency | Reasoning / call | Output (suite) |
|---|---|---|---|---|---|
| Claude Opus 5.5 | $4.03 | 2026-10-02 | 4.8s | 33 | 1,630 |
| Claude Opus 4.8 | $4.05 | 2026-07-17 | 6.1s | 0 | — |
| Qwen3.8 Max 0902 | $4.21 | 2026-09-15 | 25.2s | 546 | 5,951 |
| Qwen3.8 Max | $4.46 | 2026-08-20 | 16.8s | 589 | 6,363 |
| Fireworks Ember-1 | $4.63 | 2026-10-02 | 4.1s | 166 | 2,512 |
| Grok 4.7 | $4.7 | 2026-10-02 | 10s | 239 | 3,131 |
| MiMo V2.6 Pro UltraSpeed | $4.8 | 2026-10-02 | 2.5s | 377 | 4,661 |
| Grok 4.6 | $4.99 | 2026-08-22 | 12.7s | 628 | 6,680 |
Claude Opus 4.8 predates our per-entry price field: its cost was derived at list price on 2026-07-17 and we hold no suite token counts for it, hence the dash. Every model here scored 9/9, so on correctness the band is flat. What separates it is time:
Scoped to these eight, Ember-1 is second fastest, 1.6 seconds behind MiMo V2.6 Pro UltraSpeed, and it writes the second-fewest output tokens among the seven with token counts, after Opus 5.5. Against the reasoning-heavy members of its band it is genuinely lean: the two Qwen3.8 Max entries spend 546 and 589 reasoning tokens per call and take 25.2 and 16.8 seconds. One oddity elsewhere in the table: Grok 4.7 consumed 11,761 input tokens on the same nine prompts, a pattern we have seen before with provider-side prompt we never wrote.
Does the token-thrift claim show up?
We cannot test Fireworks' claim directly. It is a claim about Ember-1 against Kimi K3, and Kimi K3 is not in this sweep, so we have no same-day base-model run to subtract from. What we can say is narrower and more useful to a buyer:
- Against reasoning-heavy peers, the thrift is visible. 166 reasoning tokens per call is the third-lowest in the band above, behind only the two Claude Opus models.
- Against models that barely reason on easy work, it is not enough. Opus 5.5 at 33 and Sonnet 5.5 at 31 reasoning tokens per call simply do not think much about a short function. Being leaner than a heavy reasoner still leaves Ember-1 heavier than a model that skips the reasoning.
Efficiency is relative to a baseline. Fireworks chose Kimi K3 as that baseline, which is honest for someone already paying for K3. If you are choosing from the whole market, the baseline is whichever model solves your task in the fewest billed tokens, and on this suite that was not Ember-1.
Who should pay $4.63
Be clear about what nine self-contained Python functions can and cannot show. They cannot separate a frontier model from a competent small one. Solar Mini 4 scored the same 9/9 at $0.03 per 1,000 tasks and 2 seconds, both as of 2026-10-02 — the cheapest and fastest 9/9 in our set — which makes Ember-1 about 154 times the cost ($4.63 / $0.03, both priced 2026-10-02) for an identical score. That is not a verdict that Ember-1 is overpriced; it is a statement that this suite only measures the floor.
Ember-1's case rests on harder work: per its announcement, Fireworks evaluated it on agentic coding benchmarks such as SWE-bench Verified and Terminal Bench, which is where shorter reasoning traces should matter most. If you already run Kimi K3 on Fireworks for long agent sessions, Ember-1 at the same rate card is a cheap experiment. If you are shopping on short coding tasks, Sonnet 5.5 at $1.95 (priced 2026-10-02) was cheaper and faster on ours, and our cheap coding roundup covers the low end.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Ember-1 had none. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-10-02 — not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Kimi K3 on the same day. Without a same-sweep base-model run we cannot confirm or refute the headline claim about fewer tokens. That is the comparison this review most needed and does not have.
- Long agent sessions, which is where Fireworks says the saving compounds. Nine single-turn functions cannot show it.
- Input tokenization. Ember-1 logged 1,341 input tokens on prompts Opus 5.5 counted as 914. Tokenizers differ, and we did not probe a single request to rule out added provider-side prompt.
- Cached-input pricing. Our prompts are short and uncached; long repeated contexts would change the arithmetic.
- The 1,048,576-token context. Our suite is short and text-only.
- Repeat runs. One scored attempt per task. A 60-cent gap to Opus 5.5 is a measured difference on one run, not an average.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab