DeepSeek-R1

Last updated 2026-09-25. R1 (Jan 2025 Reference)

The historic foundation model that revolutionized reinforcement learning for LLMs. Trained with Group Relative Policy Optimization (GRPO) without a critic model, emitting visible, verifiable Chain-of-Thought (CoT) reasoning tokens before producing definitive answers.

Parameters
671B Total; active 37B Active / Token
Architecture
DeepSeekMoE + Pure RL via Group Relative Policy Optimization (GRPO)
Context
128K Tokens
KV cache
2.1 KB / token (FP8 MLA)
Peak input / 1M (cache miss)
0.55
Peak output / 1M
2.19

Price source: DeepSeek Models & Pricing. $0.55 / 1M input ($0.14 cache hit) $2.19 / 1M output

Sourced benchmarks

  • MATH-500: 97.3% — Rigorous 500-question subset of high school competition mathematics.
  • AIME 2024: 79.8% — First open model to achieve parity with proprietary frontier reasoning engines.
  • Codeforces Rating: 96.3 %tile (2,029 Elo) — Simulated contest performance against global human participants.
  • MMLU (Overall): 90.8% — 57 multi-discipline subjects covering humanities, STEM, and law.

All models · API pricing · FAQ