Janus-Pro

Last updated 2026-09-25. Janus-Pro (Multimodal)

An advanced open multimodal foundation model that decouples visual understanding from visual generation. Utilizes a SigLIP-L visual encoder for deep scene understanding while employing discrete Vector Quantization (VQ) codebooks and parallel prediction heads for high-fidelity text-to-image synthesis.

Parameters
7B / 1B Distillations; active Dense 7B Backbone
Architecture
Decoupled Vision: SigLIP-L Understanding + Discrete VQ Image Generation
Context
32K Tokens
KV cache
Not listed
Peak input / 1M (cache miss)
0.3
Peak output / 1M
0.9

Price source: DeepSeek Models & Pricing.

Sourced benchmarks

  • MMBench: 85.2% — Comprehensive multimodal benchmark assessing perception and reasoning.
  • GenEval: 0.81score — Systematic evaluation of compositionality in generative image models.
  • POPE (Object Hallucination): 89.4% — Polling-based object-probing evaluation benchmark.
  • Seed-Bench 2: 81.6% — Multi-level visual comprehension and reasoning assessment.

All models · API pricing · FAQ