Model News

Harvey Tenet: The Open-Weight Saving Comes From Serving, Not From the Base

Harvey announced Tenet on 2026-08-20: its first model post-trained for legal work, built on a Kimi K3 base with Fireworks AI, and pitched as running at less than a quarter the cost of leading foundation models. The story everyone is telling is that an OpenAI-backed company reached for a Chinese open-weight model. The part nobody has checked is the cost claim's mechanism. We measured Kimi K3 on our executed Python benchmark at $6.33 per 1,000 tasks — more than Claude Opus 4.8 at $4.05, nearly four times Claude Sonnet 5 at $1.67. Renting Kimi K3 from a marketplace is not cheap. So whatever Tenet's cost advantage is, it does not come from the base model being cheap — it comes from owning the weights and therefore the serving.

Chart comparing measured cost per 1,000 tasks for Kimi K3 against four leading foundation models

Every article about Tenet leads with the same framing: an OpenAI-backed legal AI company chose a Chinese open-weight base. That is true and it is interesting. It is also not the part with a number attached that anyone has checked.

What Harvey announced

Third-party throughout, from Harvey and the launch coverage, read 2026-08-22:

All of those figures are Harvey's. We have not run Tenet — it is not on any endpoint we can call — and LAB is Harvey's own evaluation.

What the base model actually costs to rent

Here is where we can contribute something. Kimi K3 is callable, and we ran it on our executed Python benchmark. These are our numbers.

ModelScoreMeasured cost / 1k tasksLatencyList price in / out
Kimi K3 — Tenet's base9/9$6.3322.8s$3.00 / $15.00
Claude Sonnet 59/9$1.677.2s—
GPT-5.49/9$1.693.6s—
Claude Opus 4.89/9$4.056.1s—
GPT-59/9$11.0022.9s—
Renting Kimi K3 costs more than renting Claude Opus 4.8Measured cost per 1,000 tasks on our executed nine-task Python benchmark. All five scored 9/9.Claude Sonnet 5$1.67GPT-5.4$1.69Claude Opus 4.8$4.05Kimi K3 — Tenet base$6.33GPT-5$11.00One scale throughout: 49.24 px per dollar. Kimi K3 priced 2026-07-30; the others on their own measurement dates.
Open weights do not mean cheap tokens when someone else is doing the serving.

Kimi K3 lists at $3.00 in and $15.00 out per million tokens on the endpoint we called. That is frontier pricing, and on our tasks it produced a frontier-sized bill — larger than three of the four leading closed models above.

Where the saving really comes from

So the arithmetic does not work if you imagine Harvey renting Kimi K3 by the token and passing on a discount. It works if you separate two things that get conflated whenever a company adopts an open-weight base:

Neither of those is visible on a price page, and both are only available to you if you hold the weights. That is the actual argument for an open-weight base, and it is a better one than the Chinese model is cheaper — which, on the numbers above, it is not. Our note on open weights versus open source covers what the licence does and does not grant.

The benchmark claims, and who measured them

Worth being precise about attribution, because the numbers are circulating without it:

None of that means the claims are wrong. A vendor evaluating on its own domain benchmark is normal and often the only benchmark that exists for the domain. It does mean there is currently no independent measurement of Tenet at all, and a self-reported improvement over a self-selected baseline is the weakest kind of evidence, however large the percentage.

Why this matters beyond legal

Tenet is a test of a specific proposition: that a vertical company can take open weights, post-train on domain data, serve it themselves, and land somewhere a general frontier model cannot reach on price. If that works in legal it works anywhere the domain corpus is defensible.

The counter-case is in the table above. Claude Sonnet 5 scored 9/9 at $1.67 and GPT-5.4 at $1.69 on our tasks — both already below a quarter of what renting Kimi K3 costs. A vertical model has to beat the rented general model, not the base it was built from. Whether Tenet does that on legal work is unmeasured, and it is the only question that matters.

How our numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price on each model's measurement date, not a billing statement. Runs go through OpenRouter. Full method on the methodology page, and prices move, so each figure carries its date.

What we did not measure

FAQ

What is Harvey Tenet? Harvey's first model post-trained for legal work, announced 2026-08-20, built on a Kimi K3 base with Fireworks AI using asynchronous reinforcement learning.

What model is Tenet based on? Kimi K3, the open-weight model from Moonshot AI.

Is Kimi K3 cheap? Not to rent. On our executed benchmark it cost $6.33 per 1,000 tasks at a $3.00 and $15.00 list price — more than Claude Opus 4.8 at $4.05.

How can Tenet cost a quarter of a foundation model then? Harvey holds the weights, so it controls the serving stack rather than paying a marketplace rate, and it says Tenet was post-trained for token efficiency. Neither lever is available to a company renting tokens.

Has anyone independently benchmarked Tenet? Not that we can point to. The published scores are Harvey's, on Harvey's LAB benchmark.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.