• Independent study
  • 240 benchmark runs
  • 8 harnesses

DeepSeek V4 HarnessChoose by outcome, speed, or cost

In an independent DeepSeek V4 Flash study, Pi had the highest reported pass rate, Claude Code was fastest, and DeepAgents had the lowest directly comparable cost per success. DeepSeek's own DSH launched two days after the study, so it was not tested.

Independent guide · Study and official release dates checked 18 August 2026.

Pi leads on reported completion, but its setup differs. Claude Code leads on median speed. DeepAgents is the clearest low-cost result. Choose the metric that matches your workload.

Benchmark facts at a glance

Harnesses
8
Tasks per harness
30
Total runs
240
Reported successes
129 (53.8%)
Highest pass rate
Pi · 66.7%*
Fastest median
Claude Code · 122.7s
Lowest comparable cost
DeepAgents · $0.045
Official DSH
Not included

Recommendation tool

Which DeepSeek V4 harness fits your goal?

Pick the outcome you value most. The answer uses the published study results and flags conditions that make a number harder to compare.

Recommendation

Highest reported pass rate: Pi

Pi completed 20 of 30 tasks (66.7%), but the study used a different reasoning setting and two model providers. OMP is the cleaner comparison when you want fewer test-condition caveats.

Published results

DeepSeek V4 harness benchmark: 8 agents, 240 runs

Composio ran 30 coding tasks per harness and reported 129 successful runs out of 240 (53.8%). Sort the table by the metric you care about.

Independent DeepSeek V4 Flash harness study published 11 August 2026
HarnessPass rateMedian timeCost / successUse this result when…
Pi66.7%132.2s$0.028*You accept the different reasoning/provider setup.
Prime62.5%*242.1s$0.131You account for 6 runs that could not be scored.
OMP56.7%272.4s$0.103You want the strongest less-qualified pass result.
Claude Code53.3%122.7s$0.195Median completion speed matters most.
Codex53.3%245.0s$0.081You want a middle-cost coding agent baseline.
DeepAgents53.3%187.1s$0.045You want the lowest directly comparable cost.
Hermes50.0%175.5s$0.056+You value low runtime-token use (~192K/task).
OpenCode46.7%129.7s$0.073You value speed and can accept the lower pass rate.

* Pi used a different reasoning configuration and two providers. Prime had 24 valid results and 6 runs that could not be scored. Hermes cost excludes an additional component reported by the study.

Read before choosing

Three details change what the DeepSeek V4 harness numbers mean

The headline ranking is useful only when its test date and setup match the decision you are making now.

01

DSH was not tested

The study was published 11 August. Official DeepSeek Harness launched 13 August, so DSH cannot be placed in this table.

02

Costs are historical

DeepSeek changed V4 pricing on 16 August. The cost-per-success values describe the study period, not today's budget.

03

Pass rate hides setup differences

Pi's leading number used a different reasoning setting and two providers. Prime's percentage uses only 24 valid runs.

Practical interpretation:

Start with Pi if you can reproduce its exact setup, OMP if you want fewer comparison caveats, Claude Code for median speed, or DeepAgents for the clearest cost result. Then rerun three tasks from your own workload.

Official option

Where official DeepSeek DSH fits

DSH is the most direct DeepSeek-owned harness, with Standard, Code, Minimal and Creator modes plus a local Web UI. The benchmark above gives no basis for calling it faster, cheaper or more accurate than the eight tested tools.

Choose DSH for first-party fit

Use it when DeepSeek's model integration, replaceable plugins and local inspection matter more than a cross-harness score.

Choose Minimal for your test

Minimal mode reduces harness assistance to persistent shell and file editing, making task comparisons easier to interpret.

Re-price every run

Use the current DeepSeek pricing calculator; do not reuse the benchmark's August 11 cost figures.

npx @deepseek-ai/dsh web

FAQ

DeepSeek V4 harness questions

DeepSeek V4 harness questions from the reference page.

  • Independent benchmark
  • DSH not tested
  • Re-run your own tasks

Pi has the highest reported pass rate in the August 11 study, but its setup differs. Claude Code is fastest by median time, while DeepAgents has the lowest directly comparable cost per successful task.

Sources

Sources and benchmark boundaries

Independent guide · Study and official release dates checked 18 August 2026.

Next step

Continue with DeepSeek

Practical pages for choosing and using agent infrastructure.