DeepSeek Harness Long-Session Cost & Compaction Recovery

Author
DeepSeekAgent.io Editorial Team
Published
Updated

Long-session cost is not determined only by the final answer. Every Agent turn can resend system prompts, history, tool definitions, and tool results. When automatic compaction does not complete in time, per-request context keeps growing and cumulative billing can greatly exceed the token count visible for one turn.

Recent community reports on 0.1.7-rc.2 and 0.2.0-rc.2 describe compaction not triggering, higher late-session cost, and interactions among Goal rounds and truncated tool results. These are reproducible cases in specific environments, not a claim that every Session is affected, but they expose a failure chain worth monitoring.

The DSH compaction path

Compaction is an optional capability outside the Agent Loop core. The automatic path evaluates two triggers before request construction:

  • pressure: estimated context approaches the configured threshold;
  • context-overflow: the provider confirms that the request exceeds its window.

The engine selects a Surface span with balanced tool calls and results, generates a summary, and replaces the old range with a new user message. compaction/start, summary, and end events record the transaction. A failed attempt remains observable instead of being represented as completed.

Why a full window can prevent recovery

Summarization itself requires a model request. If history already fills the context and the configured maximum output reserves a large region, the provider can reject the summary before compaction can save the Session.

With an 80K context and 32K maximum output, the safely usable history is far below 80K. Waiting for a meter to reach 100% is too late; the system needs room for the summary request, tool declarations, and the next output.

How truncated tool results amplify loops

Before summarization, DSH can deterministically prune oversized tool results while preserving their beginning and end. If evidence the Agent needs is removed from the middle, it may reread files, rerun commands, or repeat verification. An unfinished Goal can then continue automatic rounds and multiply the cumulative request volume.

Investigate more than “the model reasoned too much”:

  • the first turn containing […] or a spill file;
  • repeated reads of the same path or command;
  • compaction starts without useful summaries;
  • a Goal driver continuing without new progress;
  • Sub-agents copying large context concurrently.

Monitoring signals

MetricWarning sign
Input tokens per turnRapid growth across consecutive turns
Cache-miss inputFrequent changes to schemas, prompts, or prefixes
Tool repetitionThe same resources are read repeatedly
Compaction stateRepeated failures, weak shrinkage, or a stale lock
Goal roundsAutomatic continuation without new progress
Cost per minuteContinued growth disconnected from result quality

Safer configuration and recovery

  1. Set thresholds below the model's hard limit and reserve summary/output room.
  2. Give long tasks independent turn, time, and cost budgets.
  3. Read large files by range instead of reloading them every turn.
  4. Watch pruning markers and preserve critical evidence in a short note.
  5. If automatic compaction fails, stop the Agent, temporarily enlarge the window, run one manual compaction, and restore the normal setting.
  6. Verify possible side effects before resuming or retrying operations.

A reproducible evaluation

Hold the model, Profile, task, and tools constant. Compare average input tokens, cache hits, compaction count, tool calls, and cost over the first and last ten turns. Repeat with a more conservative pressure threshold. Cost differences are meaningful only when both runs meet the same task-success criteria.

Related reading

References