Space Bunny Alpha: Two Runs, No Valid Score (and the Same 502 Twice)
Space Bunny Alpha gave a wrong answer to our flatten task, and then, with the identical prompt at temperature 0, a right one. We ran the free stealth model twice on 2026-10-02. Run 1 scored 5 of the nine tasks: 4 passed, flatten was wrong. Run 2 scored 7: 6 passed, flatten passed, and parse_csv_line was wrong. Two tasks never came back in either run — roman_to_int returned provider_unavailable, HTTP 502, in both runs, and lcs_len failed at the API layer in both runs as well. So Space Bunny Alpha is excluded from our tables and has no valid score. That is not a dodge. An endpoint that answers differently twice and refuses the same task twice is telling you something more useful than a number would.
Most of what we publish is a score. This page is about a model that would not hold still long enough to get one, and about the discipline of not printing a number we do not have.
The two runs
Both runs used the same nine prompts, the same harness and the same settings, on the same day.
| Run 1 | Run 2 | |
|---|---|---|
| Date | 2026-10-02 | 2026-10-02 |
| Tasks scored | 5 of 9 | 7 of 9 |
| Passed | 4 | 6 |
| Wrong answer | flatten | parse_csv_line |
| flatten | wrong | passed |
| roman_to_int | provider_unavailable (502) | provider_unavailable (502) |
| lcs_len | API-layer failure | API-layer failure |
| Other unscored | top_k_words, parse_csv_line | none |
| Tasks with no scored answer | 4 (9 − 5) | 2 (9 − 7) |
| Verdict | EXCLUDED — incomplete in both runs | |
The tempting move is to merge the runs: take the best answer for each task and call the result a score. We do not do that for any model, because a merged best-of-two is a different benchmark from the one everything else in our tables took, and it still would not cover roman_to_int.
The answer that changed at temperature 0
flatten was wrong in run 1 and right in run 2. Same prompt, temperature 0, one attempt each. If you assume temperature 0 means the same input produces the same output, this result says the assumption does not hold here.
It should not have been a surprising assumption to break. Temperature 0 narrows sampling; it does not promise identical output, and we have written before about why it is never fully deterministic. Batching, floating-point order and routing all sit underneath it. On a stealth endpoint you also cannot see which of those changed between two calls, or whether the thing serving run 2 was configured the same way as the thing serving run 1.
What it tells you in practice: a single pass on this endpoint is not a reproducible measurement. Our one-attempt-per-task rule works for most models because most models give us the same answer twice. This one did not, on at least one task, and we only know because the first run was incomplete enough to make us run it again.
The task it would not serve, twice
roman_to_int came back as provider_unavailable with HTTP 502 in both runs. OpenRouter's API guide for this model (third-party, read 2026-10-02) describes a 502 as an upstream generation failure, meaning the request reached the provider and the provider did not produce a response.
One 502 is an outage. The same 502 on the same task in two separate runs looks less like bad luck and more like something about that particular prompt failing upstream, though we cannot see upstream and will not pretend to know what it is. When we read the endpoint page on 2026-10-02 it showed a healthy uptime figure, which is the point: an uptime number averages over everyone's traffic and says nothing about whether your request is one the endpoint will not complete.
roman_to_int was not alone. lcs_len also failed at the API layer in both runs; our records list it as an API failure without the status code, so we do not claim it was a 502. Our harness records an API-layer failure separately from a wrong answer, so both are logged as unscored rather than failed. That distinction exists because we once got it wrong: a content filter made a strong model look like it had failed five tasks when it had simply returned empty responses. Scoring a 502 as a wrong answer would have handed Space Bunny Alpha a clean-looking score that was false in both runs.
Why there is no score
Our data stores this entry as 4/5 EXCLUDED. The 4/5 is run 1's partial tally and it is not a score: it is not comparable to any 9/9, it is not a percentage you can quote, and it does not appear in our rankings. The rule is the same one that excluded models that could not finish inside a token budget. A run that does not score all nine tasks does not get a number.
The same applies to the secondary figures. The stored entry records a 4.4s mean latency, 0 reasoning tokens per call, and 1,107 input and 1,586 output tokens. Those were measured over an incomplete run, so they are not comparable to a full-suite model either.
Cost is excluded for a second reason. The endpoint is free — priced at $0 in and $0 out per million tokens on 2026-10-02 — so its derived cost is $0 per 1,000 tasks. Free models are barred from our cost rankings. A promotional zero says nothing about what the model will cost once it has a name and a rate card, and prices move even when they are not zero. For the record, the cheapest and fastest model to score 9/9 as of 2026-10-02 is Solar Mini 4, at $0.03 per 1,000 tasks priced on 2026-10-02 and a 2s mean. Space Bunny Alpha is not in that comparison at all.
One partial result we can report without a score: parse_csv_line, the task run 2 got wrong, was also missed in this sweep by Codestral 2508, Qwen3 Coder Plus, Qwen3 Coder Flash and GLM-5.3 Prime. When the Jev Router probe in the same sweep handled it, it routed that one task to Claude Opus 5.5, and it accounted for 56% of that probe's summed cost. Missing parse_csv_line puts a model in ordinary company. It is the flip on flatten and the repeated 502 that are unusual.
What we will not guess
OpenRouter's model catalogue (third-party, read 2026-10-02) describes Space Bunny Alpha as an anonymous model from a provider that has chosen not to be named, with text, image and video input, a 1,000,000-token context window, and free pricing. The same catalogue entry lists reasoning as mandatory, with a default effort of max. Our usage payloads reported 0 reasoning tokens per call.
That mismatch is exactly the kind of clue people use to name a stealth model, and we are not going to use it. We have done this before and been wrong. In August we argued from measurement that Ox Alpha was not a GLM model, largely because it reported 0 reasoning tokens. It was GLM-5.3-Flash. Under its own name, on the identical tasks, it reported 1,212. Behaviour is set by the endpoint, not by the weights: thinking budgets, reasoning visibility, system prompts and sampling are all chosen by whoever runs the endpoint, and a lab shipping anonymously has every reason to run it unusually.
So the 0 reasoning tokens could mean the reasoning is hidden from the usage field, or not happening, or something we have not thought of. It does not tell you whose model this is. Neither does anything else on this page.
What an unstable stealth endpoint can tell you
Less than a score, but not nothing.
- It can write correct code. 6 passes in a single run, on executed hidden asserts, is a real signal that the model is competent at this kind of work. It is a floor, not a ranking.
- It is not reproducible yet. A flipped answer at temperature 0 means any evaluation you run on it needs repeats, and any result you get today may not hold tomorrow.
- It has at least one input it will not serve. If your workload depends on every request completing, a stealth endpoint is a poor place to build, regardless of how good the successful answers are.
- It tells you nothing about the released model. The Ox Alpha case showed the same weights can behave very differently once the name and the serving configuration change. Measure it again when it ships.
If what you actually want is a free model to build on, our free LLM API guide covers tiers that are meant to stay up. A stealth preview is a test drive with no guarantee the car will be there next week.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task per run. An API-layer failure such as a 502 is recorded separately from a wrong answer, which is why roman_to_int and lcs_len are unscored rather than failed. Cost is derived from measured token counts at the list price captured on the run date, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
Be clear about what this suite can do even when a model completes it. Nine self-contained Python functions cannot separate a frontier model from a competent small one. They measure whether a model clears a bar, and what it costs to clear it. For a stealth endpoint that did neither, they measured something else: stability.
What we did not measure
- A third run. We stopped at two. A third might have completed, which would have produced a score we still would not fully trust given the flatten flip.
- Why the other tasks went unscored. Our records name them: roman_to_int and lcs_len in both runs, plus top_k_words and parse_csv_line in run 1. They record a 502 for roman_to_int only; for the rest we have the API-failure flag, not the status code.
- What the 502 is. We saw the status code, not the upstream cause. The same goes for lcs_len, where we do not have the code either.
- Reasoning effort. We sent no reasoning setting and took the endpoint default. We did not test whether a different effort changes the score, the tokens or the 502.
- Image and video input, the 1,000,000-token context, and tool use. Our prompts are short and text-only.
- Determinism beyond flatten. One flip on one task proves the endpoint is not deterministic; it does not tell us how often it happens.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab