The Information Machine

DeepSeek V4-Flash Launches as Sub-$0.30 Agent Model Amid Benchmark Skepticism

open · v1 · 2026-08-03 · 97 items

What

DeepSeek launched the official V4-Flash API (build 0731) in public beta on July 31, 2026, positioning it as a sub-$0.30 model optimized for agentic and coding workloads.[1][2] The model uses a 284–304B parameter Mixture-of-Experts architecture with roughly 13B active parameters per request, a 1M-token context window, and tiered reasoning modes, at $0.14 per million input tokens and $0.28 per million output tokens.[3][4] DeepSeek claims the 0731 build achieves agent benchmark scores that surpass its own V4-Pro-Preview — a result of post-training improvements applied to the same architecture without expanding model size.[5][7] SemiAnalysis publicly questioned those claims, and independent observers note the pricing advantage is clearer than the benchmark story.[5][9]

Why it matters

V4-Flash's price point — roughly 10–100x cheaper than leading frontier models — changes the economics of high-volume agentic workflows, making previously cost-prohibitive multi-step agent loops viable at scale.[3] The core unresolved question is whether the benchmark gains reflect real-world agent reliability or are artifacts of self-reported evaluation.

Open questions

  • Do V4-Flash's agent benchmark scores (Terminal-Bench 82.7, DeepSWE 54.4) hold up under third-party evaluation on real-world coding tasks, or are they optimized for DeepSeek's own benchmarks?[5][6]

  • Can the post-training-only improvement — which reportedly lifted agent scores 7x on the same 13B-active-parameter base — be reproduced or explained by independent researchers?[7]

  • Will the sub-$0.30 pricing sustain at scale, or will DeepSeek adjust as demand rises following rapid platform adoption?[15][3]

  • How will Anthropic (Claude Code) and OpenAI (Codex) respond to a model that is being directly marketed as cheaper and competitive on their primary coding-agent use case?[14][16]

Narrative

On July 31, 2026, DeepSeek released the official version of V4-Flash (tagged 0731) as a public beta API, completing what had been a preview period for the model.[1][2] The architecture is a Mixture-of-Experts design with 284–304B total parameters but only approximately 13B active per forward pass, allowing low inference cost at the chosen price: $0.14 per million input tokens and $0.28 per million output tokens.[3][4] The model supports a 1-million-token context window, up to 384K output tokens, and variable reasoning intensity (low, high, max modes). It natively supports the Responses API format and is described by DeepSeek as adapted for Codex.[5]

DeepSeek's official release claims the 0731 build has 'massively upgraded' agent capabilities, with benchmark scores 'far surpassing' its V4-Pro-Preview on internal evaluations — specifically 82.7 on Terminal-Bench 2.1 and 54.4 on DeepSWE.[5][6][3] The capability improvement came from post-training alone: the architecture and active parameter count are unchanged from the earlier preview.[7][8] Artificial Analysis raised V4-Flash's Intelligence Index score by 10 points to 50 following the update.[3] Observers note that on DeepSeek's own benchmarks, V4-Flash closes most of the gap to Opus 4.8 (Terminal 82.7 vs. 85), though the comparison relies on DeepSeek's own benchmark suite.[6]

Skepticism centers on benchmark provenance. SemiAnalysis responded to the launch announcement with the pointed question 'Please answer honestly and accurately,' signaling doubt about self-reported gains without producing a direct refutation.[5] The Neuron's analysis was more specific: the pricing advantage is unambiguous, but leaderboard performance may not translate to real-world reliability.[3] A common framing emerging from practitioners is that V4-Flash is 'not the benchmark winner' but 'the clear economic winner for coding agents,' particularly for high-volume workflows where the cost difference with Anthropic or OpenAI frontier models is an order of magnitude or more.[9]

Ecosystem adoption following the launch was rapid. Ollama added V4-Flash-0731 to its cloud on August 1, reporting a 2x speed increase within a day of listing.[10][11] OpenCode integrated it as a free built-in model.[12] Developers shared configurations using V4-Flash for execution paired with V4-Pro for planning, seeking a cost-quality balance in multi-agent pipelines.[13] Some commentators framed the launch as competitive pressure specifically on Claude Code and Codex, given V4-Flash's native Responses API support and explicit marketing around coding agent benchmarks.[14][5]

Timeline

  • 2026-07-31: DeepSeek officially launches V4-Flash API (build 0731) in public beta, claiming agent benchmark scores surpass V4-Pro-Preview. [1][2][5]
  • 2026-07-31: V4-Flash confirmed at $0.14/M input, $0.28/M output; model uses 284–304B MoE parameters with ~13B active, 1M context, 384K max output. [4][3]
  • 2026-07-31: SemiAnalysis tweets 'Please answer honestly and accurately' in response to DeepSeek's benchmark claims, signaling public skepticism. [5]
  • 2026-08-01: Ollama adds DeepSeek-V4-Flash-0731 to its cloud, noting substantial enhancements to agentic capabilities. [10]
  • 2026-08-01: OpenCode adds V4-Flash as a free built-in model in its open-source coding terminal. [12]
  • 2026-08-01: Practitioners begin pairing V4-Flash (execution) with V4-Pro (planning) in multi-agent pipelines to optimize cost-quality tradeoffs. [13]
  • 2026-08-02: The Neuron publishes analysis noting V4-Flash's post-training-only improvement lifted scores without changing architecture or active parameter count. [3][7]
  • 2026-08-02: Commentators frame V4-Flash as competitive pressure specifically targeting Claude Code and Codex given native Responses API support. [14][16]
  • 2026-08-03: Ollama cloud reports V4-Flash running 2x faster than at launch, one day after listing. [11]

Perspectives

DeepSeek (official)

Claims V4-Flash 0731 has 'massively upgraded' agent capabilities with benchmark scores far surpassing V4-Pro-Preview; emphasizes native Responses API and Codex adaptation.

Evolution: Consistent promotional framing; the 0731 build is the official release of a model previously in preview.

SemiAnalysis

Skeptical of DeepSeek's self-reported benchmark gains; publicly questioned their accuracy without presenting a direct counter-measurement.

Evolution: First stated position on this specific launch; consistent with their broader scrutiny of Chinese AI lab claims.

The Neuron (Grant Harvey)

Enthusiastic about pricing as a force that makes large-scale agent workflows economically viable; cautiously skeptical that leaderboard scores translate to real-world reliability.

Evolution: First detailed analytical take; frames the story around cost economics rather than benchmark competition.

Practitioner community (Aakash Gupta, Dorian Diaconu)

Notes the architectural approach — improving agent performance through post-training alone without expanding model size — as the more interesting technical story than raw benchmark numbers.

Evolution: Consistent across observers; emerging consensus framing from builders using the model.

Alphin Tom

Argues V4-Flash is 'not the benchmark winner' but 'may be the clear economic winner for coding agents,' separating benchmark performance from practical deployment value.

Evolution: Reflects a widely shared practitioner view separating cost from benchmark standing.

Rajath Gowda

Frames V4-Flash's launch and pricing as directly bad news for Claude Code and Codex, given the model's native Responses API support and explicit agent benchmark marketing.

Evolution: Competitive framing; consistent with community commentary positioning DeepSeek as a direct challenger.

Ollama

Adopted V4-Flash-0731 immediately on launch, reporting it 'substantially enhances agentic capabilities' and achieving 2x speed improvement within one day.

Evolution: First stated position; reflects infrastructure provider enthusiasm for distribution.

Tensions

  • DeepSeek claims V4-Flash benchmark scores 'far surpass' V4-Pro-Preview; SemiAnalysis publicly questions whether those claims are accurate. [5]
  • V4-Flash is described by practitioners as 'not the benchmark winner' but 'the clear economic winner,' while DeepSeek's own marketing emphasizes benchmark superiority over its Pro model. [9][5]
  • The post-training-only improvement claim (7x agent score gain, same architecture) is widely cited by enthusiasts as the key insight, but has not been independently verified. [7][5]
  • DeepSeek's benchmark suite is self-administered; independent evaluators like Artificial Analysis raised V4-Flash's score only 10 points, and no third-party head-to-head on real coding agent tasks has been published. [3][6]

Status: active and growing

Sources

  1. [1] 🚀 @deepseek_ai's DeepSeek-V4-Flash Official API is now LIVE in public beta. — reactive:deepseek-v4-flash-launch (2026-07-31)
  2. [2] 🚨 DeepSeek Ships V4-Flash Official API — reactive:deepseek-v4-flash-launch (2026-07-31)
  3. [3] 😺DeepSeek’s new 28-cent agent model — The Neuron (2026-08-02)
  4. [4] DeepSeek V4 Flash-0731 resmi rilis: 304B parameter MoE, context 1 juta token, output max 384K, reasoning low/high/max, d... — reactive:deepseek-v4-flash-launch (2026-07-31)
  5. [5] Please answer honestly and accurately https://t.co/OZeSaggGsu https://t.co/uFUSBhyVMl — SemiAnalysis Twitter (2026-07-31)
  6. [6] DS Flash V4 (0731) closes much of the agent gap vs Opus 4.8 on DeepSeek’s benches (Terminal 82.7 vs 85, Agents’ Last Exa... — reactive:deepseek-v4-flash-launch (2026-08-01)
  7. [7] DeepSeek reran post-training on their small model and its agent scores jumped 7x. Same architecture, same 13B active par... — reactive:deepseek-v4-flash-launch (2026-07-31)
  8. [8] DeepSeek V4 Flash 0731 is more interesting than another giant-model launch: same architecture and size, but much stronge... — reactive:deepseek-v4-flash-launch (2026-08-01)
  9. [9] DeepSeek V4 Flash is not the benchmark winner. But it may be the clear economic winner for coding agents. — reactive:deepseek-v4-flash-launch (2026-07-31)
  10. [10] DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabi... — reactive:deepseek-v4-flash-launch (2026-08-01)
  11. [11] Ollama put DeepSeek-V4-Flash on their cloud yesterday. By today it was already 2x faster. — reactive:deepseek-v4-flash-launch (2026-08-03)
  12. [12] [OpenCode just added DeepSeek V4 Flash as a FREE built-in model in its open source coding terminal 🤯 — reactive:deepseek-v4-flash-launch (2026-08-02)
  13. [13] Thats why I love using Reasonix + Deepseek V4 Flash - execution + V4 Pro - Planner to get ultimate value for money. — reactive:deepseek-v4-flash-launch (2026-07-29)
  14. [14] Bad news for Claude Code and Codex. — reactive:deepseek-v4-flash-launch (2026-08-02)
  15. [15] DeepSeek V4-Flash remains cheaper at $0.14/$0.28 and strong for pure volume. Luna’s new $0.20/$1.20 undercuts V4-Pro ($0... — reactive:deepseek-v4-flash-launch (2026-07-30)
  16. [16] 🔴 DeepSeek V4 Flash undercuts Anthropic Claude by 99%, triggering AI price war — reactive:deepseek-v4-flash-launch (2026-08-01)
  17. [17] Bad news for Claude Code and Codex. — reactive:deepseek-v4-flash-launch (2026-08-02)