DeepSeek-V4-Pro-0813

Last updated 2026-09-25. V4-Pro (Aug 2026)

DeepSeek-V4-Pro is the 1.6T-parameter Mixture-of-Experts model (49B activated) with a 1,000,000-token context, published on Hugging Face as DeepSeek-V4-Pro. The API serves checkpoint DeepSeek-V4-Pro-0813 as model name deepseek-v4-pro. The V4.1 model card's comparison table reports this checkpoint's benchmark scores; a separate V4.1-Pro had not launched when the 10 September 2026 API note was published.

Parameters
1.6T Total; active 49B Active / Token
Architecture
Massive Sparse MoE + DSpark Speculative Decoding + Persistent Context Tree
Context
1,000,000 Tokens (1M Native)
KV cache
1.2 KB / token (FP8 MLA)
Peak input / 1M (cache miss)
1.32
Peak output / 1M
3.96

Price source: DeepSeek Models & Pricing. Peak cache miss $1.32 / off-peak $0.66 per 1M input. Peak cache hit $0.044 / off-peak $0.022. Peak $3.96 / off-peak $1.98 per 1M output.

Sourced benchmarks

  • Codeforces (Rating): 3,348 — Max reasoning effort. This is not the 3,206 figure previously shown here. Source
  • GPQA Diamond (Pass@1): 92.4% — DS-V4-Pro column, maximum reasoning effort. Source
  • Terminal-Bench 2.1 (Pass@1): 87.9% — Not the 89.2 figure previously shown on this site. Source

All models · API pricing · FAQ