Benchmarks

Space Bunny Alpha: Two Runs, No Valid Score (and the Same 502 Twice)

Space Bunny Alpha gave a wrong answer to our flatten task, and then, with the identical prompt at temperature 0, a right one. We ran the free stealth model twice on 2026-10-02. Run 1 scored 5 of the nine tasks: 4 passed, flatten was wrong. Run 2 scored 7: 6 passed, flatten passed, and parse_csv_line was wrong. Two tasks never came back in either run — roman_to_int returned provider_unavailable, HTTP 502, in both runs, and lcs_len failed at the API layer in both runs as well. So Space Bunny Alpha is excluded from our tables and has no valid score. That is not a dodge. An endpoint that answers differently twice and refuses the same task twice is telling you something more useful than a number would.

DataLLM Lab article cover: Space Bunny Alpha: Two Runs, No Valid Score (and the Same 502 Twice)

Most of what we publish is a score. This page is about a model that would not hold still long enough to get one, and about the discipline of not printing a number we do not have.

The two runs

Both runs used the same nine prompts, the same harness and the same settings, on the same day.

Run 1Run 2
Date2026-10-022026-10-02
Tasks scored5 of 97 of 9
Passed46
Wrong answerflattenparse_csv_line
flattenwrongpassed
roman_to_intprovider_unavailable (502)provider_unavailable (502)
lcs_lenAPI-layer failureAPI-layer failure
Other unscoredtop_k_words, parse_csv_linenone
Tasks with no scored answer4 (9 − 5)2 (9 − 7)
VerdictEXCLUDED — incomplete in both runs
Same prompts, same settings, two different outcomesSpace Bunny Alpha, nine executed Python tasks, two runs on 2026-10-02. Each bar is the full suite.Run 1wrong: flatten4 passedwrong4 not scored5 scoredRun 2wrong: parse_csv_line6 passedwrong2 not scored7 scoredBlue = passed. Red = wrong answer. Grey = no scored answer, including roman_to_int (502) and lcs_len in both runs.One scale throughout: 60 px per task, so each full bar is 9 × 60 = 540 px. Segments: 240/60/240 and 360/60/120.
Neither bar is full, and the red segment moved. That is the whole case for exclusion.

The tempting move is to merge the runs: take the best answer for each task and call the result a score. We do not do that for any model, because a merged best-of-two is a different benchmark from the one everything else in our tables took, and it still would not cover roman_to_int.

The answer that changed at temperature 0

flatten was wrong in run 1 and right in run 2. Same prompt, temperature 0, one attempt each. If you assume temperature 0 means the same input produces the same output, this result says the assumption does not hold here.

It should not have been a surprising assumption to break. Temperature 0 narrows sampling; it does not promise identical output, and we have written before about why it is never fully deterministic. Batching, floating-point order and routing all sit underneath it. On a stealth endpoint you also cannot see which of those changed between two calls, or whether the thing serving run 2 was configured the same way as the thing serving run 1.

What it tells you in practice: a single pass on this endpoint is not a reproducible measurement. Our one-attempt-per-task rule works for most models because most models give us the same answer twice. This one did not, on at least one task, and we only know because the first run was incomplete enough to make us run it again.

The task it would not serve, twice

roman_to_int came back as provider_unavailable with HTTP 502 in both runs. OpenRouter's API guide for this model (third-party, read 2026-10-02) describes a 502 as an upstream generation failure, meaning the request reached the provider and the provider did not produce a response.

One 502 is an outage. The same 502 on the same task in two separate runs looks less like bad luck and more like something about that particular prompt failing upstream, though we cannot see upstream and will not pretend to know what it is. When we read the endpoint page on 2026-10-02 it showed a healthy uptime figure, which is the point: an uptime number averages over everyone's traffic and says nothing about whether your request is one the endpoint will not complete.

roman_to_int was not alone. lcs_len also failed at the API layer in both runs; our records list it as an API failure without the status code, so we do not claim it was a 502. Our harness records an API-layer failure separately from a wrong answer, so both are logged as unscored rather than failed. That distinction exists because we once got it wrong: a content filter made a strong model look like it had failed five tasks when it had simply returned empty responses. Scoring a 502 as a wrong answer would have handed Space Bunny Alpha a clean-looking score that was false in both runs.

Why there is no score

Our data stores this entry as 4/5 EXCLUDED. The 4/5 is run 1's partial tally and it is not a score: it is not comparable to any 9/9, it is not a percentage you can quote, and it does not appear in our rankings. The rule is the same one that excluded models that could not finish inside a token budget. A run that does not score all nine tasks does not get a number.

The same applies to the secondary figures. The stored entry records a 4.4s mean latency, 0 reasoning tokens per call, and 1,107 input and 1,586 output tokens. Those were measured over an incomplete run, so they are not comparable to a full-suite model either.

Cost is excluded for a second reason. The endpoint is free — priced at $0 in and $0 out per million tokens on 2026-10-02 — so its derived cost is $0 per 1,000 tasks. Free models are barred from our cost rankings. A promotional zero says nothing about what the model will cost once it has a name and a rate card, and prices move even when they are not zero. For the record, the cheapest and fastest model to score 9/9 as of 2026-10-02 is Solar Mini 4, at $0.03 per 1,000 tasks priced on 2026-10-02 and a 2s mean. Space Bunny Alpha is not in that comparison at all.

One partial result we can report without a score: parse_csv_line, the task run 2 got wrong, was also missed in this sweep by Codestral 2508, Qwen3 Coder Plus, Qwen3 Coder Flash and GLM-5.3 Prime. When the Jev Router probe in the same sweep handled it, it routed that one task to Claude Opus 5.5, and it accounted for 56% of that probe's summed cost. Missing parse_csv_line puts a model in ordinary company. It is the flip on flatten and the repeated 502 that are unusual.

What we will not guess

OpenRouter's model catalogue (third-party, read 2026-10-02) describes Space Bunny Alpha as an anonymous model from a provider that has chosen not to be named, with text, image and video input, a 1,000,000-token context window, and free pricing. The same catalogue entry lists reasoning as mandatory, with a default effort of max. Our usage payloads reported 0 reasoning tokens per call.

That mismatch is exactly the kind of clue people use to name a stealth model, and we are not going to use it. We have done this before and been wrong. In August we argued from measurement that Ox Alpha was not a GLM model, largely because it reported 0 reasoning tokens. It was GLM-5.3-Flash. Under its own name, on the identical tasks, it reported 1,212. Behaviour is set by the endpoint, not by the weights: thinking budgets, reasoning visibility, system prompts and sampling are all chosen by whoever runs the endpoint, and a lab shipping anonymously has every reason to run it unusually.

So the 0 reasoning tokens could mean the reasoning is hidden from the usage field, or not happening, or something we have not thought of. It does not tell you whose model this is. Neither does anything else on this page.

What an unstable stealth endpoint can tell you

Less than a score, but not nothing.

If what you actually want is a free model to build on, our free LLM API guide covers tiers that are meant to stay up. A stealth preview is a test drive with no guarantee the car will be there next week.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task per run. An API-layer failure such as a 502 is recorded separately from a wrong answer, which is why roman_to_int and lcs_len are unscored rather than failed. Cost is derived from measured token counts at the list price captured on the run date, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

Be clear about what this suite can do even when a model completes it. Nine self-contained Python functions cannot separate a frontier model from a competent small one. They measure whether a model clears a bar, and what it costs to clear it. For a stealth endpoint that did neither, they measured something else: stability.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.