Agent Harness Components: Tools, Compression & Subagents

Author
DeepSeekAgent.io Editorial Team
Published
Updated

Agent Harness products usually ship tools, planning, skills, context compression, and subagents together, making it difficult to identify which layer produces the gain. Beyond the Model: Demystifying Harness Effects in Software Engineering Agents first compares mini-SWE-agent with OpenCode, then constructs a modular NanoHarness and ablates five components with Qwen3.7-Max and DeepSeek V4 Pro.

The results point to a practical principle: harness gains come from directed, structured interaction—not simply from adding steps, tokens, or more agents.

Single-component results with DeepSeek V4 Pro

ConfigurationProgramBench scoreChange from baseline
mini-SWE-agent baseline45.14—
Structured tool registry48.68+3.54
Explicit planningSmall gain—
Lazy skillsModerate, model-dependent gain—
Context compression41.29-3.85
Task-specific subagents49.60+4.46
General subagentsBelow baseline—
Full NanoHarness51.35+6.21

The tool registry and task-specific subagents are the two most stable positive components. The full combination does not mechanically add every individual effect, but it raises DeepSeek V4 Pro from 45.14 to 51.35, close to OpenCode at 52.48 and Claude Code at 52.76 in the same experiment.

Why context compression lowers the score

For DeepSeek V4 Pro, compression cuts average prompt tokens from 10,065,204 to 2,312,701—a 77.02% reduction. The score falls by 3.85 points, while steps and tool calls decline by 61.91% and 68.87%.

ProgramBench requires an agent to preserve interface constraints, behavioral requirements, and debugging state over a long trajectory. Early summaries can remove details that remain necessary later. Compression saves substantial cost, but it is not lossless; thresholds, retention, and summary content must match the task.

This maps directly to DSH Compaction. Earlier triggering can prevent overflow, but minimizing context should not be the only objective. Long-running tasks need summaries that preserve requirements, failure evidence, and incomplete checks.

Why task-specific subagents help and general subagents hurt

Task-specific subagents increase DeepSeek V4 Pro tool calls by 73.10%, reduce prompt tokens by 9.54%, and improve the score by 4.46 points. Delegation remains aligned with concrete repository exploration, implementation, or verification goals.

General subagents increase tool calls by 82.65% and prompt tokens by 51.86%, yet score below the baseline. Without a clear responsibility and delivery contract, delegation creates duplicate exploration, context copying, and coordination overhead.

Useful multi-agent design therefore depends on boundaries:

  1. Assign each member a verifiable sub-goal.
  2. Specify inputs, outputs, and stopping conditions.
  3. Avoid having several agents read and summarize the same files.
  4. Return conclusions, evidence locations, and unresolved issues to the lead.
  5. Keep small steps in the root agent when delegation adds no independent value.

More tools do not automatically make a stronger harness

The stable positive component is a structured tool registry, not an unlimited tool catalog. A model needs to know when a tool applies, how its parameters work, what failure looks like, and how its output feeds the next turn.

The full NanoHarness raises DeepSeek V4 Pro tool calls by 83.58%, but prompt tokens by only 7.26%, while improving the score by 6.21 points. An effective harness directs additional interaction toward the task instead of letting context and calls expand together.

Direct implications for DSH Profiles

  • Minimal: suitable for short tasks and models that already operate tools reliably.
  • Standard: supplies more complete tools and continuation behavior, but needs monitoring for no-progress loops.
  • Compaction: optimize for task continuity rather than the smallest token count.
  • Subagents: prefer task-specific roles and limit generic delegation and context duplication.
  • Skills: load procedural knowledge on demand instead of repeating all instructions in every system prompt.

A useful DSH experiment should record completion, tool calls, prompt tokens, compaction events, child sessions, and failure categories—not just whether a plugin was enabled.

Related reading

References