Harvey Tenet: The Open-Weight Saving Comes From Serving, Not From the Base
Harvey announced Tenet on 2026-08-20: its first model post-trained for legal work, built on a Kimi K3 base with Fireworks AI, and pitched as running at less than a quarter the cost of leading foundation models. The story everyone is telling is that an OpenAI-backed company reached for a Chinese open-weight model. The part nobody has checked is the cost claim's mechanism. We measured Kimi K3 on our executed Python benchmark at $6.33 per 1,000 tasks — more than Claude Opus 4.8 at $4.05, nearly four times Claude Sonnet 5 at $1.67. Renting Kimi K3 from a marketplace is not cheap. So whatever Tenet's cost advantage is, it does not come from the base model being cheap — it comes from owning the weights and therefore the serving.
Every article about Tenet leads with the same framing: an OpenAI-backed legal AI company chose a Chinese open-weight base. That is true and it is interesting. It is also not the part with a number attached that anyone has checked.
What Harvey announced
Third-party throughout, from Harvey and the launch coverage, read 2026-08-22:
- Tenet is Harvey's first model post-trained for legal work, announced 2026-08-20.
- Base model: Kimi K3, the open-weight model from Moonshot AI. Post-training was done with Fireworks AI.
- Method: asynchronous reinforcement learning on a corpus of publicly available legal data, synthetic data, and human expert data simulating long-horizon legal work — contract review, diligence, evidence sweeps, litigation tasks that run for hours rather than seconds.
- Claimed results: an 82% increase in all-pass rate on LAB and 22% on LAB Contracts, relative to the Kimi K3 base. Harvey states Tenet is state of the art on LAB Contracts and second on LAB.
- Claimed economics: optimised for token efficiency, running at less than a quarter the cost of leading foundation models.
All of those figures are Harvey's. We have not run Tenet — it is not on any endpoint we can call — and LAB is Harvey's own evaluation.
What the base model actually costs to rent
Here is where we can contribute something. Kimi K3 is callable, and we ran it on our executed Python benchmark. These are our numbers.
| Model | Score | Measured cost / 1k tasks | Latency | List price in / out |
|---|---|---|---|---|
| Kimi K3 — Tenet's base | 9/9 | $6.33 | 22.8s | $3.00 / $15.00 |
| Claude Sonnet 5 | 9/9 | $1.67 | 7.2s | — |
| GPT-5.4 | 9/9 | $1.69 | 3.6s | — |
| Claude Opus 4.8 | 9/9 | $4.05 | 6.1s | — |
| GPT-5 | 9/9 | $11.00 | 22.9s | — |
Kimi K3 lists at $3.00 in and $15.00 out per million tokens on the endpoint we called. That is frontier pricing, and on our tasks it produced a frontier-sized bill — larger than three of the four leading closed models above.
Where the saving really comes from
So the arithmetic does not work if you imagine Harvey renting Kimi K3 by the token and passing on a discount. It works if you separate two things that get conflated whenever a company adopts an open-weight base:
- The weights are free. The inference is not. Open weights remove the licence fee and the vendor margin. They do not remove the GPUs. What they buy you is the right to control the serving stack — which is exactly why Fireworks AI is in this announcement, and why the cost claim is plausible even though the marketplace price is not cheap.
- Post-training for token efficiency is a cost lever in its own right. Harvey says Tenet is optimised for it. A model that reaches the same answer in fewer tokens costs less at any price per token — the mechanism we measured this week on Qwen3.8-Max, which raised its list price 36% and still cut the bill 27% by emitting 48% fewer output tokens.
Neither of those is visible on a price page, and both are only available to you if you hold the weights. That is the actual argument for an open-weight base, and it is a better one than the Chinese model is cheaper — which, on the numbers above, it is not. Our note on open weights versus open source covers what the licence does and does not grant.
The benchmark claims, and who measured them
Worth being precise about attribution, because the numbers are circulating without it:
- +82% all-pass on LAB, +22% on LAB Contracts versus the Kimi K3 base — Harvey's figures, on Harvey's benchmark.
- State of the art on LAB Contracts, second on LAB — Harvey's ranking on Harvey's benchmark.
- Under a quarter the cost of leading foundation models — Harvey's claim, with no stated methodology, no named comparison models and no date that we could find.
None of that means the claims are wrong. A vendor evaluating on its own domain benchmark is normal and often the only benchmark that exists for the domain. It does mean there is currently no independent measurement of Tenet at all, and a self-reported improvement over a self-selected baseline is the weakest kind of evidence, however large the percentage.
Why this matters beyond legal
Tenet is a test of a specific proposition: that a vertical company can take open weights, post-train on domain data, serve it themselves, and land somewhere a general frontier model cannot reach on price. If that works in legal it works anywhere the domain corpus is defensible.
The counter-case is in the table above. Claude Sonnet 5 scored 9/9 at $1.67 and GPT-5.4 at $1.69 on our tasks — both already below a quarter of what renting Kimi K3 costs. A vertical model has to beat the rented general model, not the base it was built from. Whether Tenet does that on legal work is unmeasured, and it is the only question that matters.
How our numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price on each model's measurement date, not a billing statement. Runs go through OpenRouter. Full method on the methodology page, and prices move, so each figure carries its date.
What we did not measure
- Tenet itself. It is not callable on any endpoint we have. Every Tenet figure on this page is Harvey's.
- Legal work. Our benchmark is nine Python functions. It says nothing about contract review, and a coding score does not transfer to a legal one.
- Self-hosted Kimi K3 economics. We measured the rented price. What Harvey and Fireworks pay to serve it is not public and is the whole basis of the cost claim.
- LAB. We have not seen the benchmark, its tasks, or its scoring.
FAQ
What is Harvey Tenet? Harvey's first model post-trained for legal work, announced 2026-08-20, built on a Kimi K3 base with Fireworks AI using asynchronous reinforcement learning.
What model is Tenet based on? Kimi K3, the open-weight model from Moonshot AI.
Is Kimi K3 cheap? Not to rent. On our executed benchmark it cost $6.33 per 1,000 tasks at a $3.00 and $15.00 list price — more than Claude Opus 4.8 at $4.05.
How can Tenet cost a quarter of a foundation model then? Harvey holds the weights, so it controls the serving stack rather than paying a marketplace rate, and it says Tenet was post-trained for token efficiency. Neither lever is available to a company renting tokens.
Has anyone independently benchmarked Tenet? Not that we can point to. The published scores are Harvey's, on Harvey's LAB benchmark.
DataLLM Lab