GPT-6 Luna Review: Astra's Score at a Hundredth of the Rate Card
GPT-6 Luna scored 9 out of 9 on our executed Python benchmark at $0.16 per 1,000 tasks (priced 2026-10-02) and a 4.9-second mean. The fact worth stopping on is its sibling. GPT-6 Luna Pro read 17,806 input tokens across the same nine prompts that plain Luna answered with 604 — 29.5 times the input for questions we wrote once (17,806 ÷ 604 = 29.5) — and finished at $0.49 for the identical score. The budget GPT-6 tier inherits the same Pro-tier pattern we found on GPT-6 Astra; it just costs cents instead of dollars.
GPT-6 Luna is the budget tier of OpenAI's GPT-6 line, alongside Sol and Astra. We did not test OpenAI's own pricing claims for it. We tested what Luna costs to get nine working functions out of it.
The result
| Metric | GPT-6 Luna | GPT-6 Luna Pro |
|---|---|---|
| Score | 9/9 | 9/9 |
| Measured cost / 1,000 tasks | $0.16 | $0.49 |
| Mean latency | 4.9s | 7.3s |
| Reasoning tokens per call | 201 | 370 |
| Input tokens across the suite | 604 | 17,806 |
| Output tokens across the suite | 2,763 | 5,248 |
| Priced at, in / out per 1M | $0.1 / $0.5 | $0.1 / $0.5 |
| Context window | 1,050,000 | 1,050,000 |
| Rank among 75 models at 9/9 | 9 cheapest · 23 fastest | 16 cheapest · 40 fastest |
Both runs date from 2026-10-02, and both costs were derived at the price captured that day. Plain Luna is a clean result: ninth cheapest of the 75 models that have cleared the suite, 23rd fastest of the 75, and correct on every task at temperature 0.
One number in that column is not what we expected from a budget tier. Luna spends 201 reasoning tokens per call. GPT-6 Sol spends 85 and GPT-6 Astra spends 47 on the same prompts. The cheapest GPT-6 base tier reasons the most.
Luna Pro and the 17,806-token input
We sent both tiers the identical nine prompts. Luna billed 604 input tokens; Luna Pro billed 17,806 (17,806 ÷ 604 = 29.5). The prompts did not change, so the difference is input the endpoint added before ours. Every GPT-6 Pro tier we have measured behaves the same way:
| GPT-6 tier | Input tokens, base → Pro | Cost / 1k, base → Pro | Pro multiple | Priced on |
|---|---|---|---|---|
| Luna | 604 → 17,806 | $0.16 → $0.49 | 0.49 ÷ 0.16 = 3.1x | 2026-10-02 |
| Sol | 604 → 16,898 | $2.03 → $7.45 | 7.45 ÷ 2.03 = 3.7x | 2026-10-02 |
| Astra | 604 → 17,011 | $8.19 → $35.44 | 35.44 ÷ 8.19 = 4.3x | 2026-09-15 |
Three tiers, three price points, and the Pro input lands between 16,898 and 17,806 tokens every time. That looks like one shared scaffold, not a per-model choice. On Astra Pro we confirmed the mechanism directly: a single request billed 65 prompt tokens on Astra and 1,721 on Astra Pro, 1,389 of them cached. We did not repeat that single-request probe on Luna Pro, so for Luna the overhead is inferred from suite totals, not read off one payload.
The hidden input is not the whole story on Luna. Pro also wrote 5,248 output tokens against 2,763 and reasoned 370 tokens per call against 201. At Luna's rate card an output token costs five input tokens ($0.5 ÷ $0.1), so the 2,485 extra output tokens (5,248 − 2,763) carry real weight next to the 17,202 extra input tokens (17,806 − 604). On Astra Pro the multiple was 4.3x; on Luna Pro it is 3.1x. Cheap input softens the scaffold, but it does not make it free, and it bought nothing here: same 9 out of 9, 2.4 seconds slower.
A hundredth of Astra's rate card, a fifty-first of the bill
Luna is priced at $0.1 in and $0.5 out per million (2026-10-02). Astra is priced at $10 and $50 (2026-09-15). That is exactly a hundredth on both sides of the rate card. The measured bill is not a hundredth: $8.19 ÷ $0.16 = 51x.
The gap is output. All three GPT-6 base tiers read the same 604 input tokens, so input cannot explain it. Luna wrote 2,763 output tokens across the suite; Astra wrote 1,353, and Sol 1,704. Luna produces roughly twice Astra's output (2,763 ÷ 1,353 = 2.0) for the same nine answers, which halves the discount the rate card promises. Against Sol the story repeats at a smaller scale: a twentieth of Sol's rate card ($2 ÷ $0.1), but $2.03 ÷ $0.16 = 12.7x on the bill.
None of this makes Luna expensive. It makes the rate card a poor predictor of the bill, which is a pattern we see across most of the field: verbose models erode their own discount.
Where Luna sits at the cheap end
Luna is the budget GPT-6, but it is not the budget end of the market. As of 2026-10-02, Solar Mini 4 is both the cheapest and the fastest model to score 9 out of 9 in our data, at $0.03 and 2 seconds.
| Model | Cost / 1k | Priced on | Latency | Reasoning / call | Rank, cheapest of 75 |
|---|---|---|---|---|---|
| Solar Mini 4 | $0.03 | 2026-10-02 | 2s | 0 | 1 |
| MiMo V2.6 Flash | $0.04 | 2026-10-02 | 5s | 9 | 2 |
| Ling 3.0 Flash VL | $0.07 | 2026-09-15 | 4.4s | 234 | 3 |
| DeepSeek V3.2 | $0.08 | 2026-07-30 | 7.1s | 0 | 4 |
| Qwen3 Coder Next | $0.1 | 2026-07-17 | 7s | 0 | 5 |
| Qwen3 Coder | $0.13 | 2026-10-02 | 2.7s | 0 | 8 |
| GPT-6 Luna | $0.16 | 2026-10-02 | 4.9s | 201 | 9 |
Solar Mini 4 does the same nine tasks for less than a fifth of Luna's bill ($0.16 ÷ $0.03 = 5.3x) in under half the time. MiMo V2.6 Flash matches Luna's 1,050,000-token context and its roughly five-second latency at a quarter of the cost ($0.16 ÷ $0.04 = 4x). DeepSeek V3.2's $0.08 was derived at its 2026-07-30 price; its list price has moved since, so treat that row as historical.
The claim we can make is narrower. Of the OpenAI models on our October facts sheet that scored 9 out of 9, Luna is the cheapest — below GPT-5.1-Codex-Mini at $0.45, its own Pro tier at $0.49 and GPT-5.4 mini at $0.53 (priced 2026-10-02, 2026-10-02 and 2026-07-30). The sheet does not list all 75 models; ranks 6 and 7 are not on it. So we will not call Luna the cheapest OpenAI model we have ever measured, only the cheapest OpenAI model on that sheet.
One outside signal points the same way. On 2026-10-02 we sent the same nine tasks through TypeSafe's Jev Router, which chooses a model per request; it served three of the nine with GPT-6 Luna (two_sum, lcs_len, flatten), tied with DeepSeek V4.1 Flash for the most. A router that picks a model per request reached for Luna on part of our suite; that is an observation about one router on one day, not an endorsement.
Is GPT-6 Luna worth $0.16?
Be clear about what nine self-contained Python functions can show. They cannot separate a frontier model from a competent small one: Astra, Luna and Solar Mini 4 all scored 9 out of 9. Our suite prices the floor, not the ceiling.
On that floor, Luna is a fair buy if you need to stay on OpenAI — for tooling, contracts or the 1,050,000-token context — and an unnecessary one if you do not. Against Astra it delivered the same score for a fifty-first of the bill. Against the cheap end it costs four to five times the leaders. And skip Luna Pro unless your own workload proves the scaffold earns its keep: on ours it tripled the bill and added 2.4 seconds for nothing measurable. The cheap coding roundup has the wider field.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived from measured token counts at the list price captured on the date shown beside each figure — it is not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- What Luna Pro's extra input contains, or even its per-request size. We have suite totals only; the single-request probe was run on Astra Pro, not on Luna Pro.
- Whether the Pro scaffold pays off on agentic work. It bought nothing on single-turn functions. That is our workload's blind spot, not proof it is useless.
- Why Luna reasons more than Sol or Astra. We observed 201 reasoning tokens per call against 85 and 47; we cannot see why.
- The 1,050,000-token context. Our prompts are a few hundred tokens.
- Repeat runs or load. One scored attempt per task, one day. A one-cent gap at the cheap end, such as Solar Mini 4 against MiMo V2.6 Flash, is within what a re-run could reorder.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab