DeepSeek-V3
Last updated 2026-09-25. V3 (Dec 2024 Baseline)
The baseline foundation model that established DeepSeek's global engineering leadership. Trained from scratch on 14.8T tokens for under $6M USD compute, introducing Multi-Head Latent Attention (MLA), DeepSeekMoE with auxiliary-loss-free load balancing, and native FP8 mixed precision.
- Parameters
- 671B Total; active 37B Active / Token
- Architecture
- DeepSeekMoE (256 routed + 1 shared expert) + MLA + FP8 Mixed Precision
- Context
- 128K Tokens
- KV cache
- 2.1 KB / token (FP8 MLA)
- Peak input / 1M (cache miss)
- 0.14
- Peak output / 1M
- 0.28
Price source: DeepSeek Models & Pricing. $0.14 / 1M input ($0.014 cache hit) $0.28 / 1M output
Sourced benchmarks
- MMLU (Massive Multitask): 88.5% — Standard benchmark testing factual and professional competency.
- HumanEval (Coding): 82.6% — Zero-shot standard Python programming benchmark.
- MATH-500: 90.2% — Multi-step mathematical problem solving.
- GSM8K (Grade School Math): 95.8% — Word-problem arithmetic reasoning.
All models · API pricing · FAQ