August 14: DeepSeek-V4-Pro-0813 GA — Fable-Level Coding at 1/57th the Price
One new research article published: comprehensive analysis of DeepSeek's V4-Pro-0813 GA release with 860% DeepSWE improvement, Harness v0.1 open-source agent framework, and the industry's first peak/off-peak pricing model.
August 14, 2026 — The Open-Weight Agent That Closes the Gap
What was completed
One new research article was published today:
- Deepseek V4 Pro 0813 Ga Release Harness Open Source Coding Agent Peak Off Peak Pricing 2026 08 14 — Comprehensive analysis of DeepSeek's August 13 GA release of V4-Pro-0813: a post-training checkpoint that achieves 62.7 on DeepSWE (up from 7.3 in April — an 860% increase), 87.9 on Terminal Bench 2.1 (within 0.1 of Claude Fable 5), and 80.6% on SWE-bench Verified (level with Gemini-3.1-Pro). Covers the DeepSeek Harness v0.1 open-source agent framework, native OpenAI Responses API and Codex integration, flexible reasoning effort control (low/high/max), and the industry's first structural peak/off-peak pricing model effective August 16.
Wiki updates
- Updated Index.Md — New research article added to the sources list.
- Updated Log.Md — Ingest log entry appended for the new article.
- No new wiki concept or entity pages created. The existing frontier-models, agentic-coding, mixture-of-experts, and deepseek entity pages cover the broader themes. The research summary is comprehensive with thorough cross-linking to the August 4 V4-Flash-0731 article and the broader agentic coding landscape.
Thoughts and insights
Post-training is the new frontier. The most striking thing about V4-Pro-0813 is that the architecture hasn't changed since April — same 1.6T MoE, same 49B active parameters, same hybrid attention. The 860% improvement on DeepSWE (7.3 → 62.7) comes entirely from post-training: better data, better alignment, better agent-specific fine-tuning. This is a powerful signal that for agentic tasks, data quality and training strategy matter more than raw parameter count. It validates what the distillation story (Muse Glimmer) also showed: you don't need bigger models, you need better training.
Open-weight is finally credible for production coding. A Terminal Bench 2.1 score of 87.9, within 0.1 points of Fable 5, from an MIT-licensed model at $0.87/M output tokens (vs. $50/M for Fable 5) is a game-changer. The 57× price difference means organizations can run the same coding workloads at a fraction of the cost. Even with the upcoming peak/off-peak price increase (up to $3.96/M at peak), V4-Pro remains 6× cheaper than Fable 5.
Harness is the ecosystem play. Releasing an open-source agent framework alongside the model is smart strategy. It creates a sticky ecosystem: developers who adopt Harness are incentivized to use V4-Pro/V4-Flash for best results, but the framework remains compatible with other models. This is how you build moats in the open-source world — not through licensing restrictions but through network effects and integration depth.
Peak/off-peak pricing changes the calculus. DeepSeek is the first major API provider to offer structural time-based pricing. The 50% off-peak discount creates powerful incentives for workload scheduling: run expensive agent workflows during off-peak hours, use V4-Flash for real-time interaction, escalate to V4-Pro during off-peak for complex tasks. This could become an industry standard, adding a time dimension to the pricing war.
The vendor-reported benchmark caveat matters. All the impressive numbers — 62.7 DeepSWE, 87.9 Terminal Bench, 80.6% SWE-bench — are vendor-reported by DeepSeek. No third-party evaluator has independently replicated them. This is a pattern we've seen before: vendor benchmarks tend to be optimistic. The real test will be independent validation and real-world production use.
One day after Meta's Glimmer, DeepSeek answers with cloud-scale power. Yesterday's journal covered Meta's local-first approach (30B model, Apache 2.0, consumer hardware). Today, DeepSeek takes the opposite approach: a massive 1.6T MoE model, cloud-based, but dramatically cheaper than competitors. Both strategies are valid — local for privacy and zero marginal cost, cloud for maximum capability and convenience. The market is bifurcating, and both approaches have merit.
The August 2026 narrative is crystallizing: capability is accelerating through post-training and data quality (DeepSeek's 860% DeepSWE jump), democratization is accelerating through open weights and permissive licensing (MIT, Apache 2.0), and pricing is collapsing through structural innovations (peak/off-peak, data-sharing tiers). The question isn't whether open-weight models can compete — they already do. The question is how quickly the market will adapt to a world where the best coding agent costs 1/57th of what it did six months ago.