Chinese AI lab DeepSeek shipped a public beta of its official V4-Flash API on 31 July 2026, tagged V4-Flash-0731. The release keeps the underlying Mixture-of-Experts model architecture and size unchanged from the earlier preview but bundles what DeepSeek calls a "massive upgrade" to agent capabilities, native support for OpenAI's Responses API format and full Codex compatibility — all at $0.14 per million input tokens and $0.28 per million output. Pricing puts V4-Flash roughly 2-3x cheaper than OpenAI's Luna at comparable intelligence, and the Responses-API-native path is the clearest sign yet that Chinese open-weight providers are optimising for coding agents rather than chatbots.
Codex adaptation is the headline
Most modern agent harnesses assume a backend that speaks the Responses API shape, not the older Chat Completions format. A provider that speaks Responses natively is a much closer drop-in than one that requires an adapter, which is why DeepSeek publishing configuration steps specifically for Codex integration matters. Combined with the price cut, V4-Flash-0731 is now positioned as a straight swap into existing Codex-based coding workflows.

Price and third-party benchmark
Artificial Analysis published an Intelligence, Performance and Price analysis within hours of launch, putting V4-Flash-0731 at roughly $0.03 per task at intelligence index ~50 on max reasoning effort. OpenAI Luna needs high-to-xhigh effort ($0.03-$0.04) to hit index 46-49 and max effort ($0.07) to clear index 51 — meaning Luna runs 2-3x the price of DeepSeek Flash for comparable output, though Luna is 2-5x faster per token. Throughput on OpenRouter is reported at ~93 tokens/sec. There is no multimodal support yet: V4-Flash-0731 remains text-only.
What's not in scope
DeepSeek was explicit that V4-Pro's API and its consumer App and Web products are unchanged in this release. Buyers should also confirm whether the quoted price is an off-peak baseline or a flat rate — DeepSeek's mid-July V4 launch introduced 2x peak-hour multipliers (9:00-12:00 and 14:00-18:00 Beijing time) for both Pro and Flash tiers, and Chinese-provider pricing has shifted mid-cycle before. The V4-Flash-0731 update lands alongside a busy stretch of frontier releases including Anthropic's Claude Opus 5, Google's Gemma 4 family and OpenAI's GPT-5.6 Sol launch, sharpening a market where price-to-intelligence, not raw benchmark points, is now the deciding lever for agent workloads.
Reporting based on DeepSeek's July 31 2026 announcement, DeepSeek API docs, Artificial Analysis and explainx.ai's coverage.
