DeepSeek V4 Flash 0731: Budget AI Leader Matches GPT-5.6 Luna at 60% Lower Cost
AI News

DeepSeek V4 Flash 0731: Budget AI Leader Matches GPT-5.6 Luna at 60% Lower Cost

4 min
8/1/2026
DeepSeek V4 FlashAI pricing waropen-source AIGPT-5.6 Luna

DeepSeek has dropped a bombshell in the AI pricing wars with the release of V4 Flash 0731, a major update to its budget-friendly reasoning model. The new version scores 50 on the Artificial Analysis Intelligence Index, placing it just one point behind OpenAI's GPT-5.6 Luna, while costing roughly 60% less per task. This isn't just an incremental improvement—it's a strategic repositioning that could reshape the competitive landscape for cost-sensitive AI workloads.

Intelligence Leap: From 40 to 50

The most striking improvement is the 10-point jump in the Artificial Analysis Intelligence Index, up from the original V4 Flash's score of 40. This composite benchmark evaluates models across nine rigorous evaluations, including GDPval-AA v2, Terminal-Bench v2.1, SciCode, and GPQA Diamond. The new version now trails OpenAI's budget model GPT-5.6 Luna by only a single point, a remarkable feat given the massive price differential.

DeepSeek's gains are particularly pronounced in agentic tasks. On GDPval, a benchmark simulating complex real-world office work, the model's Elo score soared from 1,189 to 1,559. It also achieved near-state-of-the-art results on Terminal-Bench 2.1 (82.7%) and DeepSWE (54.4%), showcasing its enhanced agentic capabilities. The company explicitly states it has "massively upgraded its Agent capabilities," and the model now natively supports the Responses API format and is fully adapted for Codex.

Pricing That Disrupts

The pricing is where DeepSeek V4 Flash 0731 truly shines. At $0.14 per million input tokens and $0.28 per million output tokens, it undercuts nearly every comparable model. A key driver is DeepSeek's 98% cache discount, well above the industry-standard 90%. This brings the effective cache hit price down to just $0.003 per million tokens, making repeated or similar queries extraordinarily cheap.

To put this in perspective, OpenAI's GPT-5.6 Luna, even after an 80% price cut, costs $0.20/$1.20 per million tokens. Anthropic's Claude Fable 5 (with fallback) costs $3.15 per Intelligence Index Task, while DeepSeek V4 Flash 0731 costs just $0.03. The model also uses 12% fewer tokens than its predecessor, further reducing costs. As investor Michael Burry noted, OpenAI's price cuts were "preparation for DeepSeek's new V4 models."

Architecture and Performance

Under the hood, the architecture remains unchanged: 284 billion total parameters with only 13 billion active per token, thanks to its Mixture of Experts (MoE) design. This enables remarkable efficiency, with output speeds reaching 117 tokens per second on DeepSeek's API—well above the median of 64 tokens per second for comparable open-weight models. The 1-million-token context window (roughly 1,500 A4 pages) remains a standout feature.

However, the model is text-only, supporting neither image nor multimodal inputs. It also generates output tokens very verbosely—230 million tokens were needed to evaluate it on the Intelligence Index, compared to a median of 99 million. This verbosity can inflate costs for simple tasks, though the low per-token price largely mitigates this.

continue reading below...

Open Weights, MIT License

DeepSeek continues its open-weight strategy, releasing the model under the permissive MIT license. Weights are available on Hugging Face, allowing developers to self-host or fine-tune the model. This openness contrasts with proprietary rivals like OpenAI and Anthropic, and it fuels a vibrant ecosystem of third-party providers—the model is now available through seven API providers.

The open-weight approach also enables cost optimization for high-volume users. With a blended cache hit/input/output ratio of 7:2:1, the effective price drops to just $0.06 per million tokens. For enterprises running millions of queries daily, this represents transformative savings.

Market Impact and Strategic Implications

DeepSeek's timing is impeccable. The release comes just hours after OpenAI's aggressive price cuts for its GPT-5.6 family, undercutting any temporary advantage. OpenAI Chairman Bret Taylor had argued that "token efficiency matters more" than raw model cost, but the data suggests otherwise: DeepSeek V4 Flash 0731 offers near-parity intelligence at a fraction of the price.

The model's strengths in agentic tasks position it perfectly for the growing market of AI agents and automated workflows. Combined with Moonshot's reported acquisition of 20,000 NVIDIA H200 GPUs from Alibaba, the Chinese AI ecosystem is scaling rapidly. DeepSeek's V4 Pro model is also expected soon, potentially pushing the frontier further.

The Bottom Line

DeepSeek V4 Flash 0731 is not just a budget option—it's a legitimate contender for many production workloads. Its combination of near-frontier intelligence, blazing speed, and disruptive pricing makes it the current leader in cost-performance ratio. For developers building agentic systems, coding assistants, or high-volume text processing pipelines, it's arguably the most compelling option available today.

The AI price war has entered a new phase, and DeepSeek has fired a decisive salvo. As open-weight models continue to close the intelligence gap with proprietary leaders, the competitive advantage is shifting from raw capability to cost efficiency and ecosystem openness. DeepSeek V4 Flash 0731 embodies this shift perfectly.