DeepSeek-V4.1-Flash
Last updated 2026-09-25. V4.1-Flash (Sep 2026)
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone, a Causal Encoder-Decoder that activates 8B parameters per token during prefill and 16B during decode, a 1,048,576-token context, Engram conditional memory (196B parameters), and a global KV cache of 890 bytes per token using FP4 (E2M1) caching. API model name: deepseek-flash.
- Parameters
- 552B Backbone; active 8B Prefill / 16B Decode
- Architecture
- Asymmetric Causal MoE (6 of 384 experts) + Engram (196B) + FP4 E2M1 KV Cache
- Context
- 1,048,576 Tokens (1M Native)
- KV cache
- 890 bytes / token (FP4 E2M1)
- Peak input / 1M (cache miss)
- 0.3
- Peak output / 1M
- 1.2
Price source: DeepSeek Models & Pricing. Peak cache miss $0.30 / off-peak $0.15 per 1M input. Peak cache hit $0.006 / off-peak $0.003. Peak $1.20 / off-peak $0.60 per 1M output.
Sourced benchmarks
- Codeforces (Rating): 3,471 — Instruct-model Codeforces rating at reasoning_effort=100, temperature 1.0, top_p 0.95. Source
- GPQA Diamond (Pass@1): 90.9% — GPQA Diamond pass@1 at maximum reasoning effort. Source
- Terminal-Bench 2.1 (Pass@1): 90.6% — Code-agent benchmark using the Minimal mode of DeepSeek Harness. Source
- HLE with tools (Pass@1): 63.9% — Humanity's Last Exam with tools, not the text-only number. Source
All models · API pricing · FAQ