Model–Harness fit across 66 configurations
Model–Harness Fit: 66 Configurations across Agent Tasks
- Author
- DeepSeekAgent.io Editorial Team
- Published
- Updated
Changing the Agent Harness can reverse model rankings, reshape cost, and produce different failure modes even when the model stays fixed. Finding the Right Fit: Model–Harness Interactions across Agent Tasks pairs OpenHands, DeepSeek Harness (DSH), PI, and openJiuwen with Claude Opus 5, GPT-6 Astra, GLM-5.3, Kimi K3, and DeepSeek V4 Pro across TUA-Bench, ALE-CLI, and Terminal-Bench 4. Native Codex–GPT and Claude Code–Claude runs add six reference configurations.
The study releases results for 66 configurations and 6,204 scored trajectories. Its useful question is not which harness wins overall, but how the model, harness, and task jointly determine the outcome.
Three task collections measure different workloads
| Collection | Emphasis | Tasks |
|---|---|---|
| TUA-Bench | General terminal and desktop work | 120 |
| ALE-CLI | Professional workflows and computer use | 99 |
| Terminal-Bench 4 | Hard command-line tasks | 63 |
The configurable harnesses use OpenRouter pinned to each model's first-party provider and request high reasoning effort. The study adds no skills, MCP servers, memory, or custom prompts. Context management, compaction, retries, and turn limits remain the harness defaults, so the matrix primarily measures default model–harness fit rather than the ceiling after task-specific tuning.
DeepSeek V4 Pro changes its best harness by task
| DeepSeek V4 Pro | OpenHands | DSH | PI | openJiuwen | Highest |
|---|---|---|---|---|---|
| TUA-Bench | 56.49 | 51.93 | 57.58 | 62.59 | openJiuwen |
| ALE-CLI | 51.20 | 40.94 | 47.72 | 50.83 | OpenHands |
| Terminal-Bench 4 | 3.17 | 9.52 | 6.35 | 7.94 | DSH |
DSH records the highest Terminal-Bench 4 result for DeepSeek V4 Pro, but not on the other two collections. GLM-5.3 follows the same openJiuwen → OpenHands → DSH change across tasks. Harness fit therefore has to be discussed in the context of the target workload.
Why GPT diverges under PI and DSH
On Terminal-Bench 4, GPT-6 Astra scores 60.32% with PI at $4.66 of model cost per task, versus 52.38% with DSH at $19.94. PI makes 2,560 model calls with 110 million uncached input tokens; DSH makes 11,879 calls with 673 million uncached input tokens.
This is not a general rule that fewer tools always perform better. The trajectory analysis finds that GPT sets explicit timeouts on most PI shell calls and can operate effectively with a lean scaffold. Models that emit more malformed tool calls benefit more from interception, continuation, and recovery mechanisms.
The harness must turn failure into usable feedback
The authors inspect same-task, same-model trajectories across harnesses. Models initiate 180 of 192 responses to actionable failure signals, and diagnosis or targeted repair resolves 116 of 133 such events.
What often separates success from failure is whether a signal reaches the model:
- a shell without a default timeout can turn a hung command into silence;
- stopping at an output cap removes the opportunity to continue;
- an early stuck detector can discard an otherwise viable plan;
- explicit timeout, error, and incomplete-state feedback gives the model a path to recovery.
An Agent Loop is not merely a mechanism that keeps calling a model. It is also the layer that converts environment failures into context the model can act on.
Implications for DSH evaluation
- Evaluate the model, harness, profile, and task together.
- Report cost alongside reward; more calls and longer trajectories do not guarantee more completed work.
- Before replacing the model, inspect the DSH Profile, tools, timeout behavior, continuation policy, and compaction settings.
- Retain per-task trajectories. Aggregate scores show that a gap exists; trajectories reveal whether it comes from tool use, recovery, or completion judgment.
- Choose production defaults with a representative internal task set rather than a universal leaderboard.
Relation to the earlier DSH vs PI community test
The earlier 30-task community test compares DeepSeek V4 Pro in DSH and PI while using different API paths. This paper standardizes the provider route and expands the matrix to five models, three task collections, and four configurable harnesses. They are not replications, but both show that the harness changes completion, runtime behavior, token use, and cost together.
Related reading
- Gemini Agent: A Universal Work Agent with Persistent Multi-Agent Execution
- DeepSeek Harness vs PI: 30-Task Community Benchmark
- Agent Harness Components: Tools, Compression, and Subagents
- Architecture Convergence across DSH, deepagents, and PI