DeepSeek-V4-Pro-0813
Last updated 2026-09-25. V4-Pro (Aug 2026)
DeepSeek-V4-Pro is the 1.6T-parameter Mixture-of-Experts model (49B activated) with a 1,000,000-token context, published on Hugging Face as DeepSeek-V4-Pro. The API serves checkpoint DeepSeek-V4-Pro-0813 as model name deepseek-v4-pro. The V4.1 model card's comparison table reports this checkpoint's benchmark scores; a separate V4.1-Pro had not launched when the 10 September 2026 API note was published.
- Parameters
- 1.6T Total; active 49B Active / Token
- Architecture
- Massive Sparse MoE + DSpark Speculative Decoding + Persistent Context Tree
- Context
- 1,000,000 Tokens (1M Native)
- KV cache
- 1.2 KB / token (FP8 MLA)
- Peak input / 1M (cache miss)
- 1.32
- Peak output / 1M
- 3.96
Price source: DeepSeek Models & Pricing. Peak cache miss $1.32 / off-peak $0.66 per 1M input. Peak cache hit $0.044 / off-peak $0.022. Peak $3.96 / off-peak $1.98 per 1M output.
Sourced benchmarks
All models · API pricing · FAQ