← All models
DeepSeek logo

DeepSeek V4 Flash

Planned
deepseek-ai/deepseek-v4-flash

DeepSeek V4 Flash is a 284B mixture-of-experts model activating roughly 13B parameters per token, with a 1M context window. It uses compressed latent attention, so a very long prompt costs far less memory than its length suggests. Built for work that needs a whole corpus in one pass — repository-wide code analysis, long document synthesis and multi-step reasoning over large inputs.

This model is on the roadmap but is not being served yet, so it will return model_not_found from the API. It appears here because we publish what we intend to run as openly as what we already do — the catalogue should tell you where the platform is going, not just where it is.