AI Agents

DeepSeek Harness (dsh): What Shipped, What v0.2 Adds, and Which Model to Run In It

DeepSeek Harness, command name dsh, is an MIT-licensed agent harness whose organising idea is that everything is a plugin — models, tools, skills, sessions, sandboxes, storage, the agent loop, scheduling, and the UI all mount on a kernel and can be swapped. We read the repository directly rather than quoting coverage: at 2026-10-03T08:01:58Z it had 242,563 stars and 29,079 forks, up from 183,972 when we first published on 2026-08-22, and its issue tracker is still switched off. The v0.2 line is out as release candidates, with desktop builds for macOS and Windows. None of that changes the operational point: dsh is model-agnostic, and in a harness that fires many calls per task, the model you choose moves your bill by two orders of magnitude. The last sections are our measured numbers on which models survive a loop.

Diagram of the DeepSeek Harness plugin tree showing models, tools, sandboxes, storage and UI as swappable plugins on a kernel

Update, 2026-10-03

We re-read the repository, its releases and the npm registry on 2026-10-03. What changed since the 2026-08-22 version of this page:

What actually shipped

From the project and its launch coverage, first read 2026-08-22 and re-checked 2026-10-03:

What v0.2 adds: confirmed, attributed, unsourced

A digest circulating this week says v0.2 added a desktop app for macOS and Windows, with Linux served through npm. Part of that is solid, part is older than v0.2, and part we could not trace to a source we could read. Sorted by where each claim comes from, read 2026-10-03:

ClaimTierSource
Pre-releases rc.1 (2026-09-28), rc.2 (2026-09-29), v0.2.1-alpha.1 (2026-10-03); no stable 0.2.0 yetConfirmedGitHub releases for deepseek-ai/deepseek-harness, all flagged pre-release
npm latest is 0.2.0-rc.2Confirmednpm registry for @deepseek-ai/dsh
macOS and Windows desktop builds bundle the dsh command for plugin management, with no separate Node or pnpm installConfirmedrc.2 release notes
The desktop app is new in v0.2Not quiteThe v0.1.7-rc.2 notes (2026-09-24) already add desktop onboarding and fix desktop installer packages that failed to start
v0.2 preview announced September 29, with installers for macOS on Apple silicon and 64-bit Windows from the official siteAttributedMpost, published 2026-10-02, read 2026-10-03
Linux users get it from the npm package rather than a desktop installerUnsourced, for usSearch results attribute this to DeepSeek's own post on X, which we could not load. The README does document the npm route, and the rc.2 notes mention graphical desktop launches on Linux, so Linux packaging is unsettled
Experimental Claude Code Mods compatibility layerConfirmedv0.2.1-alpha.1 notes, which say the aim is to check that the Mods API is roughly a subset of dsh plugins, not to offer full compatibility
Third-party model catalogue updated to pi-ai 0.87.1; some old model IDs removedConfirmedrc.2 release notes

The last row is the one that touches this article. If you saved a model selection in an earlier build, the rc.2 notes say it may need re-selecting — which is a moment to check what that model costs per task.

Everything is a plugin, concretely

The slogan is easy to repeat and easy to leave vague. The list of what is actually a plugin is the substance:

The plugin tree: nine categories, one kernelAnything in the blue row can be swapped without touching the rest. That includes the agent loop and the UI.Append-only session log — system prompts, reasoning, tool calls and results, subagent scheduling, every context injection · resume · fork · search · replaymodelstoolsskillssessionssandboxesstorageagentloopschedulingUICordis kernelmounts, unmounts and resolves dependencies between pluginsSource: the project’s own documentation and launch coverage, read 2026-08-22. We have not audited the code.Four runtime modes — Standard, Code, Minimal, Creator — differ only in which default plugin set they mount.
The agent loop being a plugin is the unusual part. Most harnesses treat the loop as the product.

The session log deserves its own note. It is append-only and records system prompts, reasoning, tool calls and their results, subagent scheduling, and every context injection — then supports resume, fork, search and replay. Forking a session at an arbitrary point is what makes agent behaviour debuggable rather than merely observable, and it is rare.

What the repository says, read directly

We query the GitHub API rather than quote a star count from coverage. Three snapshots:

Field2026-08-22T14:49:54Z2026-08-31T13:28:06Z2026-10-03T08:01:58Z
Stars183,972205,981242,563
Forks20,27123,87929,079
Watchers807not recorded1,042
Last push2026-08-21T12:35:08Znot recorded2026-10-03T06:02:48Z
Issue trackerdisablednot recordeddisabled
Discussionsenablednot recordedenabled
LicenceMIT—MIT

Unchanged across snapshots: language TypeScript, created 2026-08-13T11:56:32Z, topics ai-agents, cordis, dsh, dsh-plugin.

GitHub stars, read directly at three datesSame repository, same API field. Bars are drawn to one linear scale.2026-08-22183,9722026-08-31205,9812026-10-03242,563Scale: 1 px = 400 stars (0.0025 px per star); every bar starts at x = 200. Source: GitHub API, read at the timestamps in the table above.
Between the first and latest reads the count rose by 58,591 stars, and forks by 8,808.

The point that has not changed is the tracker: issues are still off on a repository with 242,563 stars, with discussions on instead. DeepSeek is publishing infrastructure, not opening a support queue, and six weeks of growth have not changed that. If you are considering building on dsh, that says more about the support you will get than any star count.

The choice dsh leaves you: which model

A model-agnostic harness moves the most expensive decision to you. These are our numbers from our nine-task Python benchmark — generated code executed against hidden asserts. Cost is derived from measured tokens at the list price on the date shown, not a bill.

ModelScoreCost / 1k tasksPriced onLatencyReasoning tokens per call
Solar Mini 49/9$0.032026-10-022s0
DeepSeek V3.29/9$0.082026-07-307.1s0
DeepSeek Chat9/9$0.102026-07-303.8s0
Qwen3 Coder Next9/9$0.102026-07-177s0
DeepSeek V4-Flash9/9$0.132026-07-1714.5s568
DeepSeek V4.1 Flash9/9$0.362026-09-1515.6s478
Claude Haiku 4.59/9$0.942026-07-293.7s0
GLM-5.19/9$4.352026-07-3023.6s1,327

Every model in that table scored 9 out of 9. As of 2026-10-03, Solar Mini 4 is both the cheapest and the fastest of the 75 models that score 9/9 on our suite. GLM-5.1 costs 54x what DeepSeek V3.2 does and 145x what Solar Mini 4 does, and the reasoning-token column is why.

DeepSeek's own newer Flash is the instructive row. V4.1 Flash costs 4.5x what V3.2 does per task, because it spends 478 reasoning tokens a call where V3.2 spends none. And its list price has moved since our run — from $0.15 / $0.6 per million tokens on 2026-09-15 to $0.3 / $1.2 on 2026-10-03 — so the $0.36 is a 2026-09-15 figure and should not be recomputed at today's price by mixing the two.

Be clear about what this table cannot tell you. Nine self-contained Python functions are a floor, not a ranking: a frontier model and a competent small one both clear them, so a 9/9 here does not separate them on the long, stateful, tool-heavy sessions a harness actually runs. What the table does separate is the price of clearing the floor.

Why reasoning tokens decide a harness bill

Reasoning tokens bill at the output rate. You cannot see them, you cannot cache them, and a harness multiplies them by every turn in a session. A dsh session that fires twenty tool calls multiplies the per-call gap twenty times over, and the per-call latency gap — 23.6 seconds for GLM-5.1 against 2s for Solar Mini 4 — compounds into wall-clock you actually feel.

That is multiplication on measured numbers, not a measurement, and it is the most useful thing to know before you point a harness at a model. We worked the same problem through for a different harness in our agent model comparison, and cutting token costs in coding agents covers the prompt-side levers. For background, what an agent harness actually is; for budgeting, agent cost. If you would rather let a router pick per call, our router measurement covers that trade.

One nuance: our tasks are single-turn. A model that reasons heavily on a one-shot function might reason less per turn inside a long session with accumulated context, or more. We have not measured that.

How our model numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived from measured token counts at the list price on each model's measurement date, not a billing statement — and prices move. Runs go through OpenRouter. Full method on the methodology page.

What we did not test

FAQ

What is DeepSeek Harness? An MIT-licensed, model-agnostic agent harness, command name dsh, released as a v0.1 developer preview on 2026-08-17. Models, tools, skills, sessions, sandboxes, storage, the agent loop, scheduling and the UI all load as plugins on the Cordis kernel.

Is DeepSeek Harness free? Yes, MIT licensed. The models you point it at are not.

Does dsh only work with DeepSeek models? No. DeepSeek, Anthropic and OpenAI via API keys, Bedrock, Vertex, Azure and Codex via native credentials, and any OpenAI-compatible endpoint, per the project documentation we read on 2026-08-22.

Is v0.2 out? As release candidates only, as of 2026-10-03: rc.1 on 2026-09-28, rc.2 on 2026-09-29, plus v0.2.1-alpha.1 on 2026-10-03. npm's latest tag is 0.2.0-rc.2.

Is there a desktop app? Yes for macOS and Windows, per the project's release notes. Linux is served by the npm route documented in the README; whether a packaged Linux desktop build is official, we could not confirm.

How many GitHub stars does DeepSeek Harness have? 242,563 at 2026-10-03T08:01:58Z, with 29,079 forks — up from 183,972 on 2026-08-22.

Which model should I run in it? On our tasks, as of 2026-10-03, Solar Mini 4 scored 9 out of 9 at $0.03 per 1,000 tasks (priced 2026-10-02) with zero reasoning tokens, and DeepSeek V3.2 did the same at $0.08 (priced 2026-07-30). Zero reasoning tokens is the profile that holds up best in a many-turn loop.

Where do I report a bug? Not the issue tracker — it is still disabled. Discussions are enabled instead.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.