DeepSeek V4 Pro vs V4.1 Flash
DeepSeek V4 Pro vs V4.1 Flash: API, Price & Agent Tests
- Author
- DeepSeekAgent.io Editorial Team
- Published
- Updated
DeepSeek currently offers V4 Pro and V4.1 Flash side by side through its official API. Both provide a one-million-token context window and up to 384K output tokens, but they target different workloads: V4.1 Flash supports image input, allows more concurrency, and costs less, while V4 Pro 0813 continues the established text-only Pro line.
API capability comparison
| Item | V4.1 Flash | V4 Pro 0813 |
|---|---|---|
| Model ID | deepseek-flash | deepseek-v4-pro |
| Text input | Supported | Supported |
| Image input | Supported | Not supported |
| Context | 1M | 1M |
| Maximum output | 384K | 384K |
| Concurrency limit | 2500 | 500 |
Tasks that need screenshots, charts, interfaces, or scanned documents should start with V4.1 Flash. Both models can serve text-based coding and Agent workloads in DeepSeek Harness, but the choice should come from the same task suite rather than the Pro or Flash label alone.
Official pricing comparison
Prices below are US dollars per million tokens. Peak and off-peak windows follow the official pricing page.
| Token type | V4.1 Flash off-peak / peak | V4 Pro off-peak / peak |
|---|---|---|
| Cache-hit input | $0.003 / $0.006 | $0.022 / $0.044 |
| Cache-miss input | $0.15 / $0.30 | $0.66 / $1.32 |
| Output | $0.60 / $1.20 | $1.98 / $3.96 |
For an Agent workload with 10 million cache-miss input tokens and 2 million output tokens, the off-peak estimate is $2.70 on V4.1 Flash and $10.56 on V4 Pro. At peak rates, the estimates are $5.40 and $21.12. Actual charges also depend on cache hits, Sub-agent fan-out, tool turns, and retries.
Agent and coding performance
DeepSeek reports strong V4.1 Flash results on Agent and coding benchmarks including Terminal Bench 2.1, DeepSWE, NL2Repo, CyberGym, and Automation-Bench; several reported results exceed those published for V4 Pro. Release evaluations may use different harnesses, tools, and settings, so these numbers indicate direction rather than guaranteeing the winner for every production task.
A better decision set includes a short repair, a cross-file change, and a long tool-driven task. Record success, manual rework, runtime, input and output tokens, cache hits, and tool turns for each model.
DeepSeek Harness guidance
V4.1 Flash is the stronger starting point when:
- the task contains images or screenshots;
- concurrent users or Agents matter;
- long-running input and output cost is important;
- the workflow should use DeepSeek's current Agent-focused model.
V4 Pro remains reasonable when:
- an established text-only production baseline already uses V4 Pro;
- migration risk matters more than immediate savings;
- internal evaluations show an advantage on a specific text workflow.
Use the explicit model ID in each Profile. Do not label deepseek-v4-pro as V4.1 Pro or assume it automatically follows future versions.
Migration checklist
- Fix the DSH version, mode, system prompt, and tool set.
- Run both
deepseek-flashanddeepseek-v4-pro. - Save the trajectory, result, tokens, runtime, and charge.
- Check vision input, structured output, and tool arguments.
- Choose against the workload's success criteria, not one benchmark score.
Related reading
- DeepSeek V4.1 Pro release status
- Configure V4.1 Flash in DeepSeek Harness
- DeepSeek Harness Sub-agent token billing