Model–Harness Fit: 66 Configurations across Agent Tasks

Author
DeepSeekAgent.io Editorial Team
Published
Updated

Changing the Agent Harness can reverse model rankings, reshape cost, and produce different failure modes even when the model stays fixed. Finding the Right Fit: Model–Harness Interactions across Agent Tasks pairs OpenHands, DeepSeek Harness (DSH), PI, and openJiuwen with Claude Opus 5, GPT-6 Astra, GLM-5.3, Kimi K3, and DeepSeek V4 Pro across TUA-Bench, ALE-CLI, and Terminal-Bench 4. Native Codex–GPT and Claude Code–Claude runs add six reference configurations.

The study releases results for 66 configurations and 6,204 scored trajectories. Its useful question is not which harness wins overall, but how the model, harness, and task jointly determine the outcome.

Three task collections measure different workloads

CollectionEmphasisTasks
TUA-BenchGeneral terminal and desktop work120
ALE-CLIProfessional workflows and computer use99
Terminal-Bench 4Hard command-line tasks63

The configurable harnesses use OpenRouter pinned to each model's first-party provider and request high reasoning effort. The study adds no skills, MCP servers, memory, or custom prompts. Context management, compaction, retries, and turn limits remain the harness defaults, so the matrix primarily measures default model–harness fit rather than the ceiling after task-specific tuning.

DeepSeek V4 Pro changes its best harness by task

DeepSeek V4 ProOpenHandsDSHPIopenJiuwenHighest
TUA-Bench56.4951.9357.5862.59openJiuwen
ALE-CLI51.2040.9447.7250.83OpenHands
Terminal-Bench 43.179.526.357.94DSH

DSH records the highest Terminal-Bench 4 result for DeepSeek V4 Pro, but not on the other two collections. GLM-5.3 follows the same openJiuwen → OpenHands → DSH change across tasks. Harness fit therefore has to be discussed in the context of the target workload.

Why GPT diverges under PI and DSH

On Terminal-Bench 4, GPT-6 Astra scores 60.32% with PI at $4.66 of model cost per task, versus 52.38% with DSH at $19.94. PI makes 2,560 model calls with 110 million uncached input tokens; DSH makes 11,879 calls with 673 million uncached input tokens.

This is not a general rule that fewer tools always perform better. The trajectory analysis finds that GPT sets explicit timeouts on most PI shell calls and can operate effectively with a lean scaffold. Models that emit more malformed tool calls benefit more from interception, continuation, and recovery mechanisms.

The harness must turn failure into usable feedback

The authors inspect same-task, same-model trajectories across harnesses. Models initiate 180 of 192 responses to actionable failure signals, and diagnosis or targeted repair resolves 116 of 133 such events.

What often separates success from failure is whether a signal reaches the model:

  • a shell without a default timeout can turn a hung command into silence;
  • stopping at an output cap removes the opportunity to continue;
  • an early stuck detector can discard an otherwise viable plan;
  • explicit timeout, error, and incomplete-state feedback gives the model a path to recovery.

An Agent Loop is not merely a mechanism that keeps calling a model. It is also the layer that converts environment failures into context the model can act on.

Implications for DSH evaluation

  1. Evaluate the model, harness, profile, and task together.
  2. Report cost alongside reward; more calls and longer trajectories do not guarantee more completed work.
  3. Before replacing the model, inspect the DSH Profile, tools, timeout behavior, continuation policy, and compaction settings.
  4. Retain per-task trajectories. Aggregate scores show that a gap exists; trajectories reveal whether it comes from tool use, recovery, or completion judgment.
  5. Choose production defaults with a representative internal task set rather than a universal leaderboard.

Relation to the earlier DSH vs PI community test

The earlier 30-task community test compares DeepSeek V4 Pro in DSH and PI while using different API paths. This paper standardizes the provider route and expands the matrix to five models, three task collections, and four configurable harnesses. They are not replications, but both show that the harness changes completion, runtime behavior, token use, and cost together.

Related reading

References