Category: DeepSeek models
This category contains 7 pages.
- DeepSeek (model family) · DeepSeek's open-weight model line, known for efficiency innovations (MLA, DeepSeekMoE) and the R1 reasoning model.
- DeepSeek LLM · DeepSeek's first general-purpose open-weight language model series (7B and 67B), released in late 2023 and accompanied by a scaling-laws study.
- DeepSeek-Coder-V2 · DeepSeek's June 2024 open-weight mixture-of-experts code model, developer-reported as competitive with closed frontier models on coding benchmarks.
- DeepSeek-R1 · DeepSeek's January 2025 open-weight reasoning model; matched o1-class results with a published RL recipe and triggered a global market repricing.
- DeepSeek-V2 · DeepSeek's May 2024 open MoE that introduced multi-head latent attention and DeepSeekMoE; started China's LLM price war.
- DeepSeek-V3 · DeepSeek's December 2024 open-weight flagship: 671B-parameter MoE reporting frontier quality from a ~$5.6M disclosed final training run.
- DeepSeek-V3.1 · DeepSeek's August 2025 hybrid open-weight model that merged thinking and non-thinking modes into a single checkpoint, succeeding both DeepSeek-V3 and DeepSeek-R1.