- Independent study
- 240 benchmark runs
- 8 harnesses
DeepSeek V4 HarnessChoose by outcome, speed, or cost
In an independent DeepSeek V4 Flash study, Pi had the highest reported pass rate, Claude Code was fastest, and DeepAgents had the lowest directly comparable cost per success. DeepSeek's own DSH launched two days after the study, so it was not tested.
Independent guide · Study and official release dates checked 18 August 2026.
Pi leads on reported completion, but its setup differs. Claude Code leads on median speed. DeepAgents is the clearest low-cost result. Choose the metric that matches your workload.
Benchmark facts at a glance
- Harnesses
- 8
- Tasks per harness
- 30
- Total runs
- 240
- Reported successes
- 129 (53.8%)
- Highest pass rate
- Pi · 66.7%*
- Fastest median
- Claude Code · 122.7s
- Lowest comparable cost
- DeepAgents · $0.045
- Official DSH
- Not included
Recommendation tool
Which DeepSeek V4 harness fits your goal?
Pick the outcome you value most. The answer uses the published study results and flags conditions that make a number harder to compare.
Recommendation
Highest reported pass rate: Pi
Pi completed 20 of 30 tasks (66.7%), but the study used a different reasoning setting and two model providers. OMP is the cleaner comparison when you want fewer test-condition caveats.
Published results
DeepSeek V4 harness benchmark: 8 agents, 240 runs
Composio ran 30 coding tasks per harness and reported 129 successful runs out of 240 (53.8%). Sort the table by the metric you care about.
| Harness | Pass rate | Median time | Cost / success | Use this result when… |
|---|---|---|---|---|
| Pi | 66.7% | 132.2s | $0.028* | You accept the different reasoning/provider setup. |
| Prime | 62.5%* | 242.1s | $0.131 | You account for 6 runs that could not be scored. |
| OMP | 56.7% | 272.4s | $0.103 | You want the strongest less-qualified pass result. |
| Claude Code | 53.3% | 122.7s | $0.195 | Median completion speed matters most. |
| Codex | 53.3% | 245.0s | $0.081 | You want a middle-cost coding agent baseline. |
| DeepAgents | 53.3% | 187.1s | $0.045 | You want the lowest directly comparable cost. |
| Hermes | 50.0% | 175.5s | $0.056+ | You value low runtime-token use (~192K/task). |
| OpenCode | 46.7% | 129.7s | $0.073 | You value speed and can accept the lower pass rate. |
* Pi used a different reasoning configuration and two providers. Prime had 24 valid results and 6 runs that could not be scored. Hermes cost excludes an additional component reported by the study.
Read before choosing
Three details change what the DeepSeek V4 harness numbers mean
The headline ranking is useful only when its test date and setup match the decision you are making now.
DSH was not tested
The study was published 11 August. Official DeepSeek Harness launched 13 August, so DSH cannot be placed in this table.
Costs are historical
DeepSeek changed V4 pricing on 16 August. The cost-per-success values describe the study period, not today's budget.
Pass rate hides setup differences
Pi's leading number used a different reasoning setting and two providers. Prime's percentage uses only 24 valid runs.
Start with Pi if you can reproduce its exact setup, OMP if you want fewer comparison caveats, Claude Code for median speed, or DeepAgents for the clearest cost result. Then rerun three tasks from your own workload.
Official option
Where official DeepSeek DSH fits
DSH is the most direct DeepSeek-owned harness, with Standard, Code, Minimal and Creator modes plus a local Web UI. The benchmark above gives no basis for calling it faster, cheaper or more accurate than the eight tested tools.
Choose DSH for first-party fit
Use it when DeepSeek's model integration, replaceable plugins and local inspection matter more than a cross-harness score.
Choose Minimal for your test
Minimal mode reduces harness assistance to persistent shell and file editing, making task comparisons easier to interpret.
Re-price every run
Use the current DeepSeek pricing calculator; do not reuse the benchmark's August 11 cost figures.
npx @deepseek-ai/dsh web
FAQ
DeepSeek V4 harness questions
DeepSeek V4 harness questions from the reference page.
- Independent benchmark
- DSH not tested
- Re-run your own tasks
Pi has the highest reported pass rate in the August 11 study, but its setup differs. Claude Code is fastest by median time, while DeepAgents has the lowest directly comparable cost per successful task.
Sources
Sources and benchmark boundaries
Independent guide · Study and official release dates checked 18 August 2026.
- Composio DeepSeek V4 Flash harness study
Published pass rate, median time, cost, token, and setup notes for eight harnesses.
- Official DeepSeek Harness repository
DSH ownership, local Web UI, modes, plugins, and release timing.
- Official DeepSeek pricing
Current model pricing used when evaluating new runs after the study.
Next step
Continue with DeepSeek
Practical pages for choosing and using agent infrastructure.
