DeepSeek-V3

Last updated 2026-09-25. V3 (Dec 2024 Baseline)

The baseline foundation model that established DeepSeek's global engineering leadership. Trained from scratch on 14.8T tokens for under $6M USD compute, introducing Multi-Head Latent Attention (MLA), DeepSeekMoE with auxiliary-loss-free load balancing, and native FP8 mixed precision.

Parameters
671B Total; active 37B Active / Token
Architecture
DeepSeekMoE (256 routed + 1 shared expert) + MLA + FP8 Mixed Precision
Context
128K Tokens
KV cache
2.1 KB / token (FP8 MLA)
Peak input / 1M (cache miss)
0.14
Peak output / 1M
0.28

Price source: DeepSeek Models & Pricing. $0.14 / 1M input ($0.014 cache hit) $0.28 / 1M output

Sourced benchmarks

  • MMLU (Massive Multitask): 88.5% — Standard benchmark testing factual and professional competency.
  • HumanEval (Coding): 82.6% — Zero-shot standard Python programming benchmark.
  • MATH-500: 90.2% — Multi-step mathematical problem solving.
  • GSM8K (Grade School Math): 95.8% — Word-problem arithmetic reasoning.

All models · API pricing · FAQ