Model Reviews

Nex N2.5 Pro: 9/9 on the Free Tier (and 66.8 Seconds Per Call)

Nex N2.5 Pro solved all nine tasks in our executed Python suite on its free endpoint — nex-agi/nex-n2.5-pro:free, 9 of 9, derived cost $0 at the list price we captured on 2026-09-15. It is also, at a 66.8-second mean, the slowest entry on the fact sheet this article draws from, and not narrowly: the next slowest thing here is Qwen3.7 Max at 25.8s, so 66.8s − 25.8s = 41.0 seconds of daylight. The interesting part is not that a free model can be correct. It is that its price of zero is the reason we will not rank it, and its latency is the reason you might not want it anyway.

DataLLM Lab article cover: Nex N2.5 Pro: 9/9 on the Free Tier (and 66.8 Seconds Per Call)

Most free endpoints we run are free because they are small, and they fail our suite in ways that are easy to write up. This one did not fail anything. It just made us wait.

The result

MetricNex N2.5 Pro (free endpoint)
Score9/9
Derived cost / 1,000 tasks$0 (priced 2026-09-15)
Mean latency66.8s
Reasoning tokens per call523
Tokens across the suite637 in / 5,436 out
List price in / out$0 / $0 per 1M (2026-09-15)
Context window262,144
Measured2026-09-16
Rank in our datanone published

That last row is not an oversight, and it is the part of this review worth reading twice. Every other 9/9 entry in our data carries a rank line — DeepSeek V4.1 Flash's reads 8 cheapest of 56 · 45 fastest. This one carries no rank line at all. We will get to why.

Set against the two ends of the field it scored the same as both of them:

Nex N2.5 Pro freeLing 3.0 Flash VLGPT-5.4 mini
Score9/99/99/9
Derived cost / 1k tasks$0 (2026-09-15)$0.07 (2026-09-15)$0.53 (2026-07-30)
Mean latency66.8s4.4s2.3s
Reasoning tokens / call5232340
Output tokens, whole suite5,4363,299969
Context262,144131,072400,000
Cost rank of 56 at 9/9barred1st cheapest10th cheapest

Do the subtraction the way a request queue does it. Choosing the free endpoint over Ling 3.0 Flash VL saves you seven cents per thousand tasks and costs 66.8s − 4.4s = 62.4 extra seconds on every single call. Against GPT-5.4 mini, the fastest 9/9 we have, it is 66.8s − 2.3s = 64.5 extra seconds. There is no workload where seven cents per thousand tasks is worth a minute of wall clock per call, and pretending otherwise is how a free tier ends up in production and then in an incident review.

Sixty-six seconds is not a token-count story

The obvious explanation for a slow model is that it writes more. That explanation does not survive contact with the sheet.

ModelOutput tokens, whole suiteReasoning tokens / callMean latency
Nex N2.5 Pro free5,43652366.8s
GLM-5.3-Flash11,8841,21224.9s
Mercury 2.59,8829892.4s
Ling 3.0 Flash VL3,2992344.4s

GLM-5.3-Flash emitted 11,884 output tokens across the suite and finished each call in 24.9s. Mercury 2.5 emitted 9,882 and finished in 2.4s. Nex N2.5 emitted 5,436 — fewer than either — and took 66.8s. It is not thinking longer than the field; 523 reasoning tokens per call is unremarkable next to GLM's 1,212. It is producing each token slower, or waiting longer before it produces the first one. From outside the endpoint we cannot tell those two apart, and we say more about that below.

Mean latency per call, six models that all scored 9/9Identical nine executed Python tasks. Mean across the nine calls, measured on each model’s own run date.Nex N2.5 Pro free66.8sQwen3.7 Max25.8sQwen3.8 Max 090225.2sGLM-5.3-Flash24.9sLing 3.0 Flash VL4.4sGPT-5.4 mini2.3sOne scale throughout: 8 px per second of mean latency.
Same score, same nine tasks. The blue bar is what free looks like on the clock.

Why a $0 model gets no rank

Our data marks this entry free, and a model marked free is barred from every cost ranking we publish. The rule is deliberate and it is not about generosity.

A derived cost of $0 is arithmetic, not economics. We compute cost the same way for every model — measured tokens multiplied by the provider's list price on the date we captured it — and when that list price is $0 in and $0 out, every token count in the world multiplies to zero. If we let that into the cost table, a promotional endpoint would sit permanently in first place ahead of a real seven-cent result, the ranking would stop being a ranking, and the number would tell you nothing about what the same model costs the day the free tier closes, tightens, or starts rate-limiting you at the worst moment. So the entry keeps its score and loses its place in the cost order. It appears in no cheapest of 56 list, and we do not print a partial rank in its place.

This is a different kind of asterisk from an exclusion, and the two get confused. An excluded entry has no valid score at all. IBM Granite 4.2 8B is the example in this batch: two of its nine tasks failed at the API layer after retries, so only seven were ever scored, and a number built on seven tasks cannot be compared to one built on nine. We publish no score for it. Nex N2.5 Pro is the opposite case — a complete, valid, nine-of-nine run whose price is the untrustworthy field.

Marked freeMarked excluded
ExampleNex N2.5 Pro freeIBM Granite 4.2 8B
Ran all nine tasks?YesNo — 2 API-layer failures
Score published9/9none
Appears in cost ranksNoNo
Reason$0 list price is a marketing figureIncomplete run

What free actually costs you

The honest case for this endpoint is narrow but real. If you are batching offline work overnight — backfilling a dataset, regenerating a corpus of small functions, anything with no human waiting — a 66.8-second mean is an inconvenience, not a cost, and 9 of 9 at $0 is genuinely hard to argue with. The 262,144-token context is twice what Ling 3.0 Flash VL carries, which is more room than most batch jobs need.

The case against it is anything interactive. A minute-plus per call rules out chat, autocomplete, agent loops that make several calls per turn, and any request sitting behind an HTTP timeout you do not control. We wrote up the levers that actually move this number in our latency guide; none of them rescue a 66.8-second mean.

And a warning that applies to the whole free shelf, not just this model: correct-and-free is rarer than the shelf suggests. In the same sweep, inference-net/schematron-v2-turbo looked ideal on the two metrics people shop on — $0.03 per 1,000 tasks priced 2026-09-15, 5.3s mean — and scored 0 of 9, missing every task in the suite. Cheap and fast is not a proxy for working. Our wider survey of what is actually free is in the free LLM API guide, and prices move often enough that any $0 should be read as a snapshot with a date on it.

One more piece of honesty about the score itself: nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models in our set clear this suite. A 9/9 here means a model can write correct, short, standard-library Python against a spec — a floor, not a ceiling. It does not tell you this model can hold a refactor together across twenty files, and nothing in this article should be read as saying it can.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why an incomplete run is excluded rather than scored low. Cost is derived from measured token counts at the provider's list price on the stated date — for this model, $0 in and $0 out captured 2026-09-15 — and is not a billing statement. Runs go through OpenRouter. The full method is on the methodology page, and the rest of the cheap end is in the cheapest-API roundup.

What we did not measure

FAQ

Is Nex N2.5 any good? On our suite, yes: 9 of 9 on the free endpoint, measured 2026-09-16. That puts it among the 56 models in our set that clear the suite, which is a floor rather than a distinction.

How much does Nex N2.5 Pro cost? The :free endpoint we ran lists at $0 in and $0 out per 1M tokens as of 2026-09-15, giving a derived cost of $0 per 1,000 tasks.

Why is it not in your cheapest list? Free pricing is barred from our cost ranks. A $0 list price makes the derived cost zero regardless of how many tokens a model burns, so ranking it would rank the promotion, not the model.

How slow is it really? 66.8 seconds mean per call — 62.4 seconds longer than Ling 3.0 Flash VL and 64.5 seconds longer than GPT-5.4 mini, both of which scored the same 9 of 9.

Should I use it? For unattended batch work where nobody is waiting, it is a reasonable free option. For anything interactive, the latency disqualifies it before the price gets a vote.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.