Category: 2025 model releases
This category contains 51 pages.
- AFM-4.5B · The first from-scratch Arcee Foundation Model: a 4.5B-parameter enterprise-focused open-weight LLM trained on 8 trillion tokens, later extended to a 64K context window.
- Claude 3.7 Sonnet · Anthropic's February 2025 hybrid reasoning model: one network serving instant and extended-thinking modes, launched alongside Claude Code.
- Claude 4 · Anthropic's May 2025 generation (Opus 4, Sonnet 4), built for agentic coding; Opus 4 was the first model deployed under ASL-3 safeguards.
- Claude Haiku 4.5 · Anthropic's October 2025 small model, the first Haiku with extended thinking, offering developer-reported near-Sonnet-4 coding performance at a fraction of the cost.
- Claude Opus 4.1 · Anthropic's August 2025 incremental upgrade to Claude Opus 4, with developer-reported gains in agentic coding and long-horizon reasoning.
- Claude Opus 4.5 · Anthropic's November 2025 flagship: developer-reported coding leadership with a major price cut, closing the Claude 4.x series.
- Claude Sonnet 4.5 · Anthropic's September 2025 mid-tier hybrid reasoning model, marketed at launch as the company's strongest model for coding and computer use.
- Comma (models) · Pair of 7-billion-parameter open-weight language models released in June 2025 by EleutherAI and Common Pile collaborators, trained solely on the openly licensed Common Pile v0.1 corpus.
- DeepSeek R1 release shock · The January 2025 market and policy reaction to DeepSeek-R1: a historic one-day selloff in AI-linked equities and a rethink of compute assumptions.
- DeepSeek-R1 · DeepSeek's January 2025 open-weight reasoning model; matched o1-class results with a published RL recipe and triggered a global market repricing.
- DeepSeek-V3.1 · DeepSeek's August 2025 hybrid open-weight model that merged thinking and non-thinking modes into a single checkpoint, succeeding both DeepSeek-V3 and DeepSeek-R1.
- Devstral · Mistral AI's open-weight agentic coding model line, launched May 2025 with All Hands AI and tuned for software-engineering agent workflows.
- ERNIE 5 · Baidu's fifth-generation ERNIE flagship series: the natively omni-modal ERNIE 5.0 (November 2025) and the efficiency-focused ERNIE 5.1 (May 2026).
- Gemini 2.5 · Google DeepMind's 2025 thinking-by-default generation; 2.5 Pro led leaderboards through spring 2025.
- Gemini 3 · Google DeepMind's November 2025 generation; launch-led major benchmarks and deepened the Deep Think reasoning tier.
- Genie · Google DeepMind's series of generative interactive world models (Genie, Genie 2, Genie 3) that produce playable environments from images or text.
- GLM-4.5 · Zhipu AI's July 2025 open-weight mixture-of-experts model, positioned as a unified agentic, reasoning, and coding foundation model.
- GLM-4.6 · Zhipu AI's late-2025 open-weight mixture-of-experts model, an agentic-coding-focused successor to GLM-4.5 with a 200K-token context window.
- GPT-4.5 · OpenAI's February 2025 research-preview chat model, positioned as its largest scaling-focused (non-reasoning) model and retired from the API in mid-2025.
- GPT-5 · OpenAI's August 2025 flagship: a unified system routing between fast and reasoning models, ending the GPT-4-era model picker.
- GPT-5.1 · OpenAI's November 2025 update to GPT-5, pairing a warmer default conversational style with adaptive reasoning in Instant and Thinking variants.
- GPT-5.2 · OpenAI's December 2025 update to the GPT-5 series, shipped in Instant, Thinking, and Pro variants amid intensified competition with Google's Gemini 3.
- gpt-oss · OpenAI's August 2025 open-weight models (120B and 20B MoE, Apache 2.0), its first open weights since GPT-2.
- Grok 3 · xAI's February 2025 flagship model, trained on the Colossus supercluster and released with a dedicated reasoning mode and the DeepSearch agent.
- Grok 4 · xAI's July 2025 flagship reasoning model, trained with large-scale reinforcement learning and released alongside a multi-agent Grok 4 Heavy variant.
- Grok 4.1 · xAI's November 2025 update to Grok 4, focused on emotional intelligence, creative writing, and reduced hallucinations, with a developer-reported first place on the LMArena text leaderboard at launch.
- IBM Granite · IBM's family of open-weight enterprise language models, spanning dense, mixture-of-experts, code, and hybrid Mamba/transformer variants.
- Kimi k1.5 · Moonshot AI's January 2025 multimodal reasoning model, trained with long-context reinforcement learning and released the same week as DeepSeek-R1.
- Kimi K2 · Moonshot AI's July 2025 trillion-parameter open-weight MoE; put open weights at the agentic frontier and introduced the MuonClip optimizer.
- Ling (models) · Ant Group's open-weight mixture-of-experts model family (Bailing line), notable for frontier-scale training partly on Chinese domestic accelerators.
- Llama 4 · Meta's April 2025 MoE generation (Scout, Maverick; Behemoth unreleased); a mixed reception that preceded Meta's Superintelligence Labs reorganization.
- Magistral · Mistral AI's first reasoning model line, released in June 2025 as the open-weight Magistral Small (24B) and the proprietary Magistral Medium.
- MAI-1 · Microsoft AI's in-house foundation model line: the mixture-of-experts MAI-1-preview entered public testing on LMArena in August 2025 and began feeding Copilot text features.
- MiniMax-M1 · MiniMax's June 2025 open-weight hybrid-attention reasoning model, developer-reported as the first large-scale open-weight model combining lightning attention with mixture-of-experts.
- MiniMax-M2 · MiniMax's October 2025 open-weight mixture-of-experts model aimed at agentic and coding workflows, released under the MIT license.
- MiniMax-Text-01 · MiniMax's January 2025 open-weight language model, notable for its hybrid lightning-attention architecture and developer-reported 4-million-token context window.
- Mistral Medium 3 · Mistral AI's May 2025 proprietary enterprise model, marketed as near-frontier performance at a substantially lower price.
- Mistral Small · Mistral AI's mid-size model line, launched February 2024 and later released open-weight, positioned for low-latency deployment below Mistral Large.
- o3 · OpenAI's frontier reasoning model, previewed in December 2024 and released in April 2025, noted for developer-reported gains on ARC-AGI, math, and coding benchmarks.
- o4-mini · OpenAI's April 2025 small reasoning model, released alongside o3 with tool use and image-based reasoning at a lower price point.
- Qwen2.5-Max · Alibaba's January 2025 proprietary large-scale mixture-of-experts flagship, positioned against DeepSeek-V3 and Western frontier models.
- Qwen2.5-VL · Alibaba's January 2025 open-weight vision-language model series (3B to 72B) with document parsing, object grounding, long-video understanding, and GUI-agent capabilities.
- Qwen3 · Alibaba's April 2025 open-weight model family introducing hybrid thinking and non-thinking modes across dense and mixture-of-experts variants.
- Qwen3-Coder · Alibaba's July 2025 open-weight agentic coding model, a 480B-parameter mixture-of-experts specialist derived from the Qwen3 family.
- Qwen3-Max · Alibaba's September 2025 proprietary flagship of the Qwen3 generation, developer-reported to exceed one trillion parameters.
- QwQ-32B · Alibaba Qwen team's open-weight 32B reasoning model, previewed in November 2024 and released in March 2025 with developer-reported parity to much larger reasoners.
- Seed-OSS · ByteDance Seed's August 2025 open-weight 36B dense language model family, notable for a long native context window and a user-controllable thinking budget.
- Seedream · ByteDance Seed's text-to-image generation model family (Seedream 2.0-4.0), noted for bilingual Chinese-English prompting and integrated image editing.
- Sora 2 · OpenAI's September 2025 video-and-audio generation model, launched alongside a consumer Sora app with likeness-based 'cameos'.
- Trinity Mini · Arcee AI's 26B-parameter sparse mixture-of-experts model (3B active) for agents and tool orchestration, released open-weight under Apache 2.0 in December 2025.
- Trinity Nano · Arcee AI's 6B-parameter sparse mixture-of-experts model (about 1B active per token) with a 128K context window, aimed at edge, embedded, and offline deployments.