DeepSeek-V4-Flash

DeepSeek-V4-Flash is the model I leave as the default. Official copy: 284 billion total parameters, 13 billion active, reasoning close to Pro, cheaper and faster, good enough on simple agent tasks. Chat’s Instant Mode is this path. The API name is deepseek-v4-flash, documented as DeepSeek-V4-Flash-0731 when I checked.

It shares the V4 contract: 1M context, 384K max output, Thinking on by default with effort mapping, JSON, tools, Responses API. FIM is non-thinking only. The pricing table listed a 2500 concurrency cap — five times Pro’s. When deepseek-chat and deepseek-reasoner were in their sunset window, they routed to Flash non-thinking and thinking. I now call Flash by its real name so logs stay readable after that sunset.

I escalate to DeepSeek-V4-Pro when the ticket involves a long agent loop or a codebase that already fought me. For screenshots, see vision-exp.

Next step

This is an independent overview page, not a DeepSeek account or checkout. To use the product, open the official site.

official DeepSeek website