Model Reviews

Fireworks Ember-1 Review (9/9, $4.63): Cheaper Per Token Than Opus 5.5, Dearer Per Task

Fireworks Ember-1 is a model whose whole pitch is spending fewer tokens, and on our executed Python benchmark it did the job: 9 out of 9 at $4.63 per 1,000 tasks (priced 2026-10-02) with a 4.1-second mean. The surprise is the comparison it loses. Ember-1 lists at $3 in / $15 out per 1M, a quarter below Claude Opus 5.5 at $4 / $20. Measured the same day on the same nine prompts, Opus 5.5 cost $4.03 and Ember-1 cost $4.63, both priced 2026-10-02 — because Ember-1 wrote 54% more output tokens. A cheaper rate card did not survive contact with the token count.

DataLLM Lab article cover: Fireworks Ember-1 Review (9/9, $4.63): Cheaper Per Token Than Opus 5.5, Dearer Per Task

Fireworks is best known as the place you run other labs' open-weight models. Ember-1 ships under its own name, as fireworks/ember-1 — though, as Fireworks itself says, it is another lab's model underneath. That makes it a useful test of a narrow question: when an inference host tunes a model for its own economics, does the saving reach your bill?

The result

MetricFireworks Ember-1
Score9/9
Measured cost / 1,000 tasks$4.63 (priced 2026-10-02)
Mean latency4.1s
Reasoning tokens per call166
Tokens across the suite1,341 in / 2,512 out
List price in / out$3 / $15 per 1M (2026-10-02)
Context window1,048,576
Measured2026-10-02

Among the 75 models in our set that score 9/9, Ember-1 ranks 53rd cheapest and 14th fastest. That is the profile of a quick, mid-premium model: near the fast end, past the midpoint on cost.

What Ember-1 is, according to Fireworks

Everything in this section is third-party — Fireworks' own claims and listings, read 2026-10-02 — not something we measured.

Two things we deliberately leave out. Several aggregator pages repeat a parameter count, but Fireworks' announcement does not state one, so neither do we. And we found no licence statement for the Ember-1 weights themselves; that Kimi K3 is open-weight does not tell you Ember-1 is. Our own launch-day look at the base model is in the Kimi K3 review.

Cheaper per token, dearer per task

Anthropic's two models new to this sweep went through the same nine prompts on the same day, which gives an unusually clean comparison:

ModelList in / out per 1MOutput tokens (suite)Reasoning / callCost / 1k tasksLatency
Fireworks Ember-1$3 / $152,512166$4.634.1s
Claude Opus 5.5$4 / $201,63033$4.034.8s
Claude Sonnet 5.5$2 / $101,56931$1.953.4s

All three scored 9/9; all prices and costs are as of 2026-10-02. Ember-1's rate card is exactly 0.75 of Opus 5.5's on both input and output ($3 / $4, $15 / $20). Yet the measured bill runs the other way: $4.63 against $4.03, 60 cents more per thousand tasks, both priced 2026-10-02. The reason is in the token columns. Ember-1 emitted 2,512 output tokens across the suite to Opus 5.5's 1,630 — 1.54 times as many — and about five times the reasoning tokens per call, 166 against 33. At $15 per million, output is the expensive side of Ember-1's card, five times its input rate, so extra output tokens land directly on the bill.

Ember-1 does buy something for that 60 cents: it was 0.7 seconds faster than Opus 5.5 at the mean. Against Claude Sonnet 5.5, though, it buys nothing on this suite. Sonnet 5.5 scored the same 9/9 at $1.95 (priced 2026-10-02), 0.7 seconds faster than Ember-1, so Ember-1 costs 2.4 times as much for a slower answer. This is the lesson of our piece on reasoning tokens deciding the bill again: the rate card is an input, the token count is the multiplier, and only the product matters.

The $4 to $5 band

To place Ember-1 among its price peers we took the eight 9/9 models in our fact sheet with a full entry and a measured cost between $4.03 and $4.99 per 1,000 tasks (capture dates in the table). This is a local set chosen by cost, not every model at that price, and each cost carries the date its list price was captured:

ModelCost / 1kPriced onLatencyReasoning / callOutput (suite)
Claude Opus 5.5$4.032026-10-024.8s331,630
Claude Opus 4.8$4.052026-07-176.1s0—
Qwen3.8 Max 0902$4.212026-09-1525.2s5465,951
Qwen3.8 Max$4.462026-08-2016.8s5896,363
Fireworks Ember-1$4.632026-10-024.1s1662,512
Grok 4.7$4.72026-10-0210s2393,131
MiMo V2.6 Pro UltraSpeed$4.82026-10-022.5s3774,661
Grok 4.6$4.992026-08-2212.7s6286,680

Claude Opus 4.8 predates our per-entry price field: its cost was derived at list price on 2026-07-17 and we hold no suite token counts for it, hence the dash. Every model here scored 9/9, so on correctness the band is flat. What separates it is time:

Same score, same price band, a tenfold spread in waitMean latency, seconds. All eight scored 9/9; label shows cost per 1,000 tasks.MiMo V2.6 Pro UltraSpeed2.5s · $4.8Fireworks Ember-14.1s · $4.63Claude Opus 5.54.8s · $4.03Claude Opus 4.86.1s · $4.05Grok 4.710s · $4.7Grok 4.612.7s · $4.99Qwen3.8 Max16.8s · $4.46Qwen3.8 Max 090225.2s · $4.21One scale throughout: 20 px per second of mean latency. Costs priced on each model’s own capture date.
Within these eight, only MiMo V2.6 Pro UltraSpeed answered faster than Ember-1.

Scoped to these eight, Ember-1 is second fastest, 1.6 seconds behind MiMo V2.6 Pro UltraSpeed, and it writes the second-fewest output tokens among the seven with token counts, after Opus 5.5. Against the reasoning-heavy members of its band it is genuinely lean: the two Qwen3.8 Max entries spend 546 and 589 reasoning tokens per call and take 25.2 and 16.8 seconds. One oddity elsewhere in the table: Grok 4.7 consumed 11,761 input tokens on the same nine prompts, a pattern we have seen before with provider-side prompt we never wrote.

Does the token-thrift claim show up?

We cannot test Fireworks' claim directly. It is a claim about Ember-1 against Kimi K3, and Kimi K3 is not in this sweep, so we have no same-day base-model run to subtract from. What we can say is narrower and more useful to a buyer:

Efficiency is relative to a baseline. Fireworks chose Kimi K3 as that baseline, which is honest for someone already paying for K3. If you are choosing from the whole market, the baseline is whichever model solves your task in the fewest billed tokens, and on this suite that was not Ember-1.

Who should pay $4.63

Be clear about what nine self-contained Python functions can and cannot show. They cannot separate a frontier model from a competent small one. Solar Mini 4 scored the same 9/9 at $0.03 per 1,000 tasks and 2 seconds, both as of 2026-10-02 — the cheapest and fastest 9/9 in our set — which makes Ember-1 about 154 times the cost ($4.63 / $0.03, both priced 2026-10-02) for an identical score. That is not a verdict that Ember-1 is overpriced; it is a statement that this suite only measures the floor.

Ember-1's case rests on harder work: per its announcement, Fireworks evaluated it on agentic coding benchmarks such as SWE-bench Verified and Terminal Bench, which is where shorter reasoning traces should matter most. If you already run Kimi K3 on Fireworks for long agent sessions, Ember-1 at the same rate card is a cheap experiment. If you are shopping on short coding tasks, Sonnet 5.5 at $1.95 (priced 2026-10-02) was cheaper and faster on ours, and our cheap coding roundup covers the low end.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Ember-1 had none. Cost is derived — measured input and output token counts multiplied by the list price captured 2026-10-02 — not a billing statement, and prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.