DeepSeek-R1
Last updated 2026-09-25. R1 (Jan 2025 Reference)
The historic foundation model that revolutionized reinforcement learning for LLMs. Trained with Group Relative Policy Optimization (GRPO) without a critic model, emitting visible, verifiable Chain-of-Thought (CoT) reasoning tokens before producing definitive answers.
- Parameters
- 671B Total; active 37B Active / Token
- Architecture
- DeepSeekMoE + Pure RL via Group Relative Policy Optimization (GRPO)
- Context
- 128K Tokens
- KV cache
- 2.1 KB / token (FP8 MLA)
- Peak input / 1M (cache miss)
- 0.55
- Peak output / 1M
- 2.19
Price source: DeepSeek Models & Pricing. $0.55 / 1M input ($0.14 cache hit) $2.19 / 1M output
Sourced benchmarks
- MATH-500: 97.3% — Rigorous 500-question subset of high school competition mathematics.
- AIME 2024: 79.8% — First open model to achieve parity with proprietary frontier reasoning engines.
- Codeforces Rating: 96.3 %tile (2,029 Elo) — Simulated contest performance against global human participants.
- MMLU (Overall): 90.8% — 57 multi-discipline subjects covering humanities, STEM, and law.
All models · API pricing · FAQ