Model Review

Kimi K2.7 Code Review: Moonshot Shipped a Leaner Specialist, Not a Smarter Model

Moonshot AI released Kimi K2.7 Code on June 12, 2026 with two claims: it cuts thinking tokens by roughly 30% versus K2.6, and it improves on K2.6 across six benchmarks. Five weeks later there is enough independent evidence to check both. The efficiency claim is not just confirmed — it is beaten. Artificial Analysis, running the identical harness on both models, found output tokens fell from 170M to 100M, a ~41% drop. The capability claim splits: on the same AA harness the general Intelligence Index fell, 44 to 42, while on DeepSWE — the one independent benchmark measuring the agentic coding this model was specialized for — K2.7 Code posts 31% against K2.6's 24%. Our own executed test adds the axis nobody else publishes: what a correct answer actually costs. This is the assembled picture: what Moonshot reported, what independent evaluators measured, and what we measured ourselves.

Kimi K2.7 Code review — vendor-reported, independent and first-party benchmark evidence side by side

What Kimi K2.7 Code is

Kimi K2.7 Code is Moonshot AI's coding-focused model, released June 12, 2026 under a Modified MIT license with open weights on Hugging Face. The specs, confirmed across Moonshot's model card, its platform docs and OpenRouter:

How this is sourced. Specs and the benchmark table are from Moonshot's Hugging Face model card and platform docs. Independent figures are from Artificial Analysis (checked live July 17, 2026) and the DeepSWE leaderboard (updated July 16, 2026) — both are re-run periodically and can move. First-party figures are from our own executed coding benchmark. For per-token rates and the cache-hit discount, see our Kimi API pricing guide.

What Moonshot published — and quietly conceded

Moonshot published a six-row table comparing K2.7 Code against K2.6, GPT-5.5 and Claude Opus 4.8. Most coverage of this model reprints that table and stops. It is worth actually reading, because the interesting part is not the improvement over K2.6 — it is what the other two columns say.

Benchmark (vendor-run)Kimi K2.6Kimi K2.7 CodeGPT-5.5Claude Opus 4.8
Kimi Code Bench v250.962.069.067.4
Program Bench48.353.669.163.8
MLS Bench Lite26.735.135.542.8
Kimi Claw 24/7 Bench42.946.952.850.4
MCP Atlas69.476.079.481.3
MCP Mark Verified72.881.192.976.4

Verbatim from Moonshot's Hugging Face model card, July 2026. Bold cells mark the best score in each row. Every number here was produced by Moonshot, in Moonshot's harness.

Read across the rows against Moonshot's own new model: GPT-5.5 beats K2.7 Code in all six. Opus 4.8 beats it in five of six. K2.7 Code wins exactly one comparison against Opus — MCP Mark Verified, 81.1 to 76.4 — and beats GPT-5.5 on nothing. Even MLS Bench Lite, the row with K2.7's biggest gain over K2.6 (+31.5% relative), still loses to both frontier models.

That is a vendor publishing a table where its new model loses eleven of twelve head-to-head comparisons on a benchmark suite it chose and ran itself. It is an unusually honest piece of marketing, and almost nobody covering this model has pointed it out.

One correction to a criticism that circulates about this table: it is not true that every benchmark is a Moonshot invention. Kimi Code Bench v2, Kimi Claw 24/7 Bench, Program Bench and MLS Bench Lite are Moonshot-created. But MCP Mark Verified is MCPMark, an independent academic benchmark of 127 expert-curated MCP tasks (arXiv 2509.24002) that Moonshot did not create. The accurate criticism is narrower and still fair: every score in this table is vendor-run and self-reported, whoever designed the benchmark.

What independent evaluators found

Here is where the story gets genuinely interesting, because the independent evidence exists — it is just scattered across an Artificial Analysis model page, its K2.6 page, the DeepSWE leaderboard and a kernel researcher's public run. Assembled, it does something none of those sources does alone: it splits Moonshot's pitch apart along a seam the launch post never mentions.

Artificial Analysis has evaluated K2.7 Code across nine public evals — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR — scoring it 42 on the Intelligence Index, ranked #6 of 97 (as of July 17, 2026). Crucially, it ran the identical index on K2.6. That gives us the exact generation-over-generation comparison Moonshot asks us to take on faith:

Same harness, both models: efficiency up, general capability slightly downArtificial Analysis Intelligence Index · K2.6 vs K2.7 Code · checked July 17, 2026Output tokens to run the indexlower is leaner170MK2.6100MK2.7 Code~41% fewerIntelligence Index scorehigher is smarter44K2.642K2.7 Codeslightly lower
Chart: DataLLM Lab — figures measured and published by Artificial Analysis, read July 17, 2026; chart and framing ours. Both panels are drawn to true scale from zero: the token drop is large, the capability drop is small but real. The AA index measures general capability across reasoning, knowledge and coding — see DeepSWE below for the agentic-coding-only picture. AA scores are live and may be revised.

So Moonshot's efficiency claim is not just corroborated — it is beaten, on somebody else's harness. K2.7 Code needed ~41% fewer output tokens than K2.6 to complete the same nine evaluations. Moonshot's careful wording was approximately 30% on average for thinking tokens; the independent number is larger.

42 versus 44. On general capability, the independent index says K2.7 Code came out slightly behind its predecessor, while Moonshot's table shows it improving in all six of its own rows. That is not a catastrophe — two points is a small gap and K2.7 Code is still #6 of 97 — but it contradicts the shape of the vendor story.

Except on the one axis this model was actually built for. DeepSWE, an independent long-horizon agentic coding benchmark of 113 tasks across 91 repositories, lists K2.7 Code at 31% (±1%) on its July 16, 2026 board — 13th of 15 models, far below GPT-5.6-sol's 73% and Claude Opus 4.8's 59%. But DeepSWE's own launch post scored K2.6 at 24% (±2%). Same benchmark, same evaluator: K2.7 Code is meaningfully better at the agentic coding it was specialized for. Carry one caveat — K2.6 is not on the current board, so this is a cross-snapshot comparison rather than a same-run one like AA's, and it deserves less weight for that reason.

A third independent signal cuts the other way. Researcher Elliot Arledge ran K2.7 Code on KernelBench-Hard, a public GPU kernel optimization benchmark, against K2.6. The MoE kernel result regressed from K2.6's 0.222 to 0.157, and two of its kernels died on its own bugs before the 45-minute clock ran out. But his read adds a nuance worth carrying: on five of six problems, K2.7 Code wrote real Triton kernels where K2.6 had leaned on library wrappers. His summary — K2.7 is more honest but not more capable — is a sharper description of this model than anything in the launch materials.

Two claims, three verdicts

Put all three evidence classes in one place and the picture resolves. To our knowledge this table does not exist anywhere else — the published coverage of this model has exactly one of these columns, never all three.

QuestionMoonshot (vendor-run)Independent (July 2026)DataLLM Lab (first-party, executed)
Thinking-token efficiency vs K2.6~30% fewer, on average170M → 100M output tokens, identical harness (~41% fewer)272 reasoning tokens/task — 2nd-leanest of 9 reasoners
General capability vs K2.6Better in all 6 published rowsWorse: AA Index 42 vs 44; KernelBench MoE 0.157 vs 0.222Not tested — we never ran K2.6
Agentic coding vs K2.6Better in all 6 published rowsBetter: DeepSWE 31% vs 24% (cross-snapshot)Not tested — we never ran K2.6
Capability vs frontierLoses all 6 rows to GPT-5.5; 5 of 6 to Opus 4.8AA Index 42 (#6 of 97); DeepSWE 31% (13th of 15)9/9 correct — tied with 9 other models
Real cost of a correct answerNot publishedNot published$1.34 per 1,000 executed tasks
SpeedNot published46.8 tok/s on Kimi's own API (#51 of 97)10.4s avg/task via OpenRouter
SWE-bench VerifiedNot publishedAbsent from Vals AI board as of 7/14/2026Not tested

Vendor figures from Moonshot's model card. Independent figures from Artificial Analysis (read 7/17/2026), DeepSWE (7/16/2026 board; K2.6 figure from DeepSWE's launch post), Vals AI (7/14/2026) and Elliot Arledge's public KernelBench-Hard run. First-party from our nine-task executed benchmark, July 2026.

The verdict splits three ways, not two. Moonshot's efficiency story holds up under independent scrutiny and then some. Its improves on K2.6 story holds on the one independent benchmark that measures what this model was specialized for — agentic coding, 24% to 31% on DeepSWE — and fails on the broader general-capability index, 44 down to 42. K2.7 Code is a leaner specialist: better at the narrow thing, slightly worse at everything else, and much cheaper to run either way.

That is not a scandal — it is what specialization looks like when a vendor is honest enough to publish a table its own model mostly loses. A coding model that thinks 41% less and gets better at coding while shedding a little general capability is a legitimately useful thing to ship. It is just not the uncomplicated across-the-board upgrade the launch implied.

Our own executed test

We ran K2.7 Code alongside 12 other models on nine coding tasks where every answer is executed against hidden tests — generate the code, run it, score pass or fail. No self-reported anything. Full method: how we test.

Our contribution to the token-efficiency question is a different axis from Artificial Analysis's. AA compared K2.7 Code against itself, one generation back. We compared it against the field:

Reasoning tokens per coding task — Kimi vs the field9 reasoning models · same 9 executed tasks · DataLLM Lab, July 2026 · lower is leanerDeepSeek V4-Pro732MiniMax M3623DeepSeek V4-Flash568GLM 5.2559Grok 4.3482StepFun 3.7 Flash450Nemotron 3 Ultra373Kimi K2.7 Code272GPT-5.5176Four other models tested — Qwen3 Coder Next, Mistral Medium 3.5, Claude Sonnet 5, Claude Opus 4.8 — emitted zero reasoning tokens and are not shown.
Chart: DataLLM Lab — measured reasoning tokens per task from our own executed nine-task run, July 2026. Kimi K2.7 Code (highlighted) is the second-leanest reasoner of the nine, behind only GPT-5.5. Four of the 13 models tested emitted no reasoning tokens at all and are excluded.

At 272 reasoning tokens per task, K2.7 Code is the second-leanest of the nine models that reason at all — behind only GPT-5.5 (176) and running at roughly 37% of DeepSeek V4-Pro's 732. This is a cross-vendor result, and it corroborates AA's within-family result from a completely different angle: the leanness shows up against its own predecessor and against the rest of the field.

Two honest caveats. First, we never tested K2.6, so we cannot verify Moonshot's 30% figure at all — a cross-vendor ranking and a within-family delta measure different things. AA tested K2.6 and got ~41%; that is the number to cite, not ours. Second, lean is relative. AA's own page calls K2.7 Code's 100M somewhat verbose against a 92M field average. Both readings are true: on our nine short executed coding tasks it is unusually terse, while across AA's broader mix of long-form reasoning evals it still writes more than the median model. Task shape decides which you see.

What a correct answer actually costs

This is the axis no independent source publishes, and it is the one that decides real deployments. Ten of our 13 models scored a perfect 9/9. Correctness is table stakes on standard coding tasks. What separates them is an 88x cost spread for identical output.

ModelScore$ / 1,000 tasksAvg latencyReasoning tokens/task
Qwen3 Coder Next9/9$0.107.0s0
DeepSeek V4-Flash9/9$0.1314.5s568
Mistral Medium 3.59/9$0.872.9s0
Nemotron 3 Ultra9/9$1.078.1s373
Kimi K2.7 Code9/9$1.3410.4s272
Claude Sonnet 59/9$1.677.2s0
GLM 5.29/9$1.9912.3s559
Claude Opus 4.89/9$4.056.1s0
GPT-5.59/9$8.8310.5s176

Selected rows from our 13-model run, July 2026. Cost = real measured token usage x list price, extrapolated to 1,000 tasks. Full field and the three models that scored below 9/9 — DeepSeek V4-Pro, Grok 4.3 and StepFun 3.7 Flash, all at 8/9: the complete coding cost benchmark.

Kimi K2.7 Code lands at $1.34 per 1,000 tasks — mid-pack. It sits between Nemotron 3 Ultra ($1.07) and Claude Sonnet 5 ($1.67), costs about a third of Claude Opus 4.8 ($4.05) and roughly a sixth of GPT-5.5 ($8.83) for the identical 9/9. It is also about 13x more expensive than Qwen3 Coder Next, which scored the same 9/9 at $0.10.

Which price list, and why it matters. K2.7 Code has two published list prices and they disagree by ~25% on input: Moonshot official is $0.19 cache-hit / $0.95 cache-miss input and $4.00 output; OpenRouter lists $0.72 / $3.50. Our $1.34 uses OpenRouter's rates, the catalog our harness prices against. At Moonshot's direct cache-miss rates the figure would be somewhat higher; with cache hits on a repeated-context agent, meaningfully lower. Stating this is the difference between a reproducible number and a decorative one.

The latency figure needs the same discipline. Our 10.4s per task is host-dependent, not a property of the model. Artificial Analysis measures 46.8 tokens/sec on Kimi's own API — #51 of 97, and low for open-weight models of its size — while third-party hosts do better. If you benchmark this model on a fast host and get a very different number, that is expected. Pick the host, then measure.

Test it against the field yourself

Kimi K2.7 Code, Qwen3 Coder Next, DeepSeek, GLM, Claude and GPT-5.5 — one OpenAI-compatible endpoint, 300+ models on a single key. Run your own nine tasks before you trust anyone's table, including ours.

The gap that remains

One thing genuinely is still missing: SWE-bench Verified. Moonshot did not publish one, and the model is absent from the Vals AI SWE-bench Verified leaderboard as of its 7/14/2026 update — five weeks post-release. For reference, that board currently reads: Claude Fable 5 95.00%, Claude Opus 4.8 88.60%, Grok 4.5 86.60%, GPT-5.5 82.60%, Claude Opus 4.7 82.00%, Gemini 3.5 Flash 78.80%. No Kimi entry.

Warning — do not trust the SWE-bench numbers you will find for this model. Several sites publish confident-looking K2.7 Code scores: 78.2% SWE-Bench, 60.4% SWE-bench Verified, 68.5% LiveCodeBench, SWE-bench Pro 58.6. These are mutually contradictory and appear on no primary source. They are almost certainly LLM-generated filler. A related trap: Vals AI's own K2.7 Code detail page renders every benchmark as 0.0% — every row, including SWE-bench, LiveCodeBench and Terminal-Bench 2.1, while simultaneously showing a 49.99% overall average. That is a page-rendering artifact, not a score. The only sound inference from Vals is the negative one: absence from the leaderboard.

So the honest framing of the gap is narrow, not sweeping. It is not true that nobody has independently tested this model — Artificial Analysis has, across nine public evals including GPQA Diamond and Terminal-Bench v2.1, and DeepSWE has it on its board at 31%. It is true that the single most-cited agentic coding benchmark still has no entry for it. Expect that to change; a K2.7 Code SWE-bench Verified result could land any week and date this section.

Operational constraints worth knowing

Before you wire this into anything, four things in Moonshot's docs will surprise you:

Who should actually use it

A decision rule, given everything above:

The broader lesson generalizes past this model. When a vendor gives you a generational-improvement claim and an efficiency claim in the same launch post, they are not equally checkable and they do not have to be equally true — and the improvement claim is rarely true uniformly. Here the efficiency claim survived independent scrutiny and got better; the improvement claim turned out to be true on the narrow axis the model was specialized for and false on the broad one. The only way to know which is which is to look for someone who ran both models on the same harness — and, where nobody has, to run it yourself.

FAQ

Is Kimi K2.7 Code better than Kimi K2.6?

Depends what you measure. Moonshot's table: better in all six rows. Artificial Analysis, identical harness on both: general Intelligence Index 42 vs K2.6's 44 — slightly worse. KernelBench-Hard: MoE kernel 0.157 vs 0.222 — worse. DeepSWE, independent agentic coding: 31% vs 24% — better. Fair summary as of July 2026 — meaningfully leaner, better at the coding it was specialized for, slightly worse at general capability. A trade, not a straight upgrade.

Has Kimi K2.7 Code been independently benchmarked?

Yes, by at least two evaluators. Artificial Analysis evaluated it across nine public evals (GPQA Diamond, Terminal-Bench v2.1, SciCode, HLE and more) at Intelligence Index 42, #6 of 97 as of 7/17/2026. DeepSWE lists it at 31% (±1%), 13th of 15, as of 7/16/2026. Elliot Arledge ran it on public KernelBench-Hard. Still missing: a SWE-bench Verified entry.

Does Kimi K2.7 Code really use 30% fewer thinking tokens?

Independent data says the direction is right and the magnitude is bigger. Artificial Analysis needed 170M output tokens for K2.6 and 100M for K2.7 Code on the identical index — ~41% fewer. We never tested K2.6 so we cannot check the 30% ourselves, but our run puts K2.7 Code at 272 reasoning tokens/task, 2nd-leanest of 9 reasoners.

What is Kimi K2.7 Code's SWE-bench Verified score?

There isn't one. Moonshot published none, and it is absent from Vals AI's SWE-bench Verified board as of 7/14/2026. Ignore the 78.2% / 60.4% / 68.5% figures circulating on content farms — mutually contradictory, no primary source, almost certainly generated filler. If you want an independent agentic coding number that does exist, use DeepSWE: 31%, 13th of 15.

How much does Kimi K2.7 Code cost to run?

List prices disagree: Moonshot official $0.19 cache-hit / $0.95 input / $4.00 output; OpenRouter $0.72 / $3.50. At OpenRouter rates our executed test measured $1.34 per 1,000 coding tasks — between Nemotron 3 Ultra ($1.07) and Claude Sonnet 5 ($1.67), ~1/6 of GPT-5.5's $8.83 for the same 9/9.

Should I use Kimi K2.7 Code for coding?

Reasonable mid-priced open-weights pick, not the value leader. It scored 9/9 — so did nine others, including Qwen3 Coder Next at 13x less. It earns its spot when you need 256K context, self-hostable Modified MIT weights, image input, and lean reasoning in long agentic loops.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.