DSH long-session cost and compaction
DeepSeek Harness Long-Session Cost & Compaction Recovery
- Author
- DeepSeekAgent.io Editorial Team
- Published
- Updated
Long-session cost is not determined only by the final answer. Every Agent turn can resend system prompts, history, tool definitions, and tool results. When automatic compaction does not complete in time, per-request context keeps growing and cumulative billing can greatly exceed the token count visible for one turn.
Recent community reports on 0.1.7-rc.2 and 0.2.0-rc.2 describe compaction not triggering, higher late-session cost, and interactions among Goal rounds and truncated tool results. These are reproducible cases in specific environments, not a claim that every Session is affected, but they expose a failure chain worth monitoring.
The DSH compaction path
Compaction is an optional capability outside the Agent Loop core. The automatic path evaluates two triggers before request construction:
pressure: estimated context approaches the configured threshold;context-overflow: the provider confirms that the request exceeds its window.
The engine selects a Surface span with balanced tool calls and results, generates a summary, and replaces the old range with a new user message. compaction/start, summary, and end events record the transaction. A failed attempt remains observable instead of being represented as completed.
Why a full window can prevent recovery
Summarization itself requires a model request. If history already fills the context and the configured maximum output reserves a large region, the provider can reject the summary before compaction can save the Session.
With an 80K context and 32K maximum output, the safely usable history is far below 80K. Waiting for a meter to reach 100% is too late; the system needs room for the summary request, tool declarations, and the next output.
How truncated tool results amplify loops
Before summarization, DSH can deterministically prune oversized tool results while preserving their beginning and end. If evidence the Agent needs is removed from the middle, it may reread files, rerun commands, or repeat verification. An unfinished Goal can then continue automatic rounds and multiply the cumulative request volume.
Investigate more than “the model reasoned too much”:
- the first turn containing
[…]or a spill file; - repeated reads of the same path or command;
- compaction starts without useful summaries;
- a Goal driver continuing without new progress;
- Sub-agents copying large context concurrently.
Monitoring signals
| Metric | Warning sign |
|---|---|
| Input tokens per turn | Rapid growth across consecutive turns |
| Cache-miss input | Frequent changes to schemas, prompts, or prefixes |
| Tool repetition | The same resources are read repeatedly |
| Compaction state | Repeated failures, weak shrinkage, or a stale lock |
| Goal rounds | Automatic continuation without new progress |
| Cost per minute | Continued growth disconnected from result quality |
Safer configuration and recovery
- Set thresholds below the model's hard limit and reserve summary/output room.
- Give long tasks independent turn, time, and cost budgets.
- Read large files by range instead of reloading them every turn.
- Watch pruning markers and preserve critical evidence in a short note.
- If automatic compaction fails, stop the Agent, temporarily enlarge the window, run one manual compaction, and restore the normal setting.
- Verify possible side effects before resuming or retrying operations.
A reproducible evaluation
Hold the model, Profile, task, and tools constant. Compare average input tokens, cache hits, compaction count, tool calls, and cost over the first and last ten turns. Repeat with a more conservative pressure threshold. Cost differences are meaningful only when both runs meet the same task-success criteria.
Related reading
- Why did a simple task take 50 turns?
- DSH Sub-agent usage and billing
- DeepSeek Harness dynamic tools and KV Cache