Category: Reasoning models
This category contains 45 pages.
- Chain-of-thought prompting · Eliciting step-by-step reasoning before the answer; discovered as a prompting trick in 2022, later internalized by trained reasoning models.
- Claude 3.7 Sonnet · Anthropic's February 2025 hybrid reasoning model: one network serving instant and extended-thinking modes, launched alongside Claude Code.
- Claude 4 · Anthropic's May 2025 generation (Opus 4, Sonnet 4), built for agentic coding; Opus 4 was the first model deployed under ASL-3 safeguards.
- Claude Fable 5 · Anthropic's June 2026 frontier model, first of the Mythos-class tier above Opus; briefly suspended under a US export-control directive.
- Claude Haiku 4.5 · Anthropic's October 2025 small model, the first Haiku with extended thinking, offering developer-reported near-Sonnet-4 coding performance at a fraction of the cost.
- Claude Opus 4.1 · Anthropic's August 2025 incremental upgrade to Claude Opus 4, with developer-reported gains in agentic coding and long-horizon reasoning.
- Claude Opus 4.5 · Anthropic's November 2025 flagship: developer-reported coding leadership with a major price cut, closing the Claude 4.x series.
- Claude Opus 4.6 · Anthropic's February 2026 flagship, adding a beta 1M-token context window, adaptive thinking, and effort controls to the Opus line.
- Claude Opus 4.8 · Anthropic's May 2026 Opus flagship, adding dynamic workflows, a discounted fast mode, and developer-reported gains in coding honesty; the last Opus release before Claude Fable 5.
- Claude Opus 5 · Anthropic's July 2026 Opus release, positioned as approaching Claude Fable 5's capability at half its price; the first Opus-tier model of the Claude 5 generation.
- Claude Sonnet 4.5 · Anthropic's September 2025 mid-tier hybrid reasoning model, marketed at launch as the company's strongest model for coding and computer use.
- Claude Sonnet 5 · Anthropic's June 2026 mid-tier model, marketed as its most agentic Sonnet, released as the default model for Free and Pro plans.
- DeepSeek-R1 · DeepSeek's January 2025 open-weight reasoning model; matched o1-class results with a published RL recipe and triggered a global market repricing.
- DeepSeek-V3.1 · DeepSeek's August 2025 hybrid open-weight model that merged thinking and non-thinking modes into a single checkpoint, succeeding both DeepSeek-V3 and DeepSeek-R1.
- Gemini 2.5 · Google DeepMind's 2025 thinking-by-default generation; 2.5 Pro led leaderboards through spring 2025.
- Gemini 3 · Google DeepMind's November 2025 generation; launch-led major benchmarks and deepened the Deep Think reasoning tier.
- Gemini 3.5 · Google DeepMind's 2026 Gemini generation: Flash shipped at I/O in May 2026 as the default Gemini model, while the Pro flagship remained unreleased as of July 2026.
- GLM-4.5 · Zhipu AI's July 2025 open-weight mixture-of-experts model, positioned as a unified agentic, reasoning, and coding foundation model.
- GLM-4.6 · Zhipu AI's late-2025 open-weight mixture-of-experts model, an agentic-coding-focused successor to GLM-4.5 with a 200K-token context window.
- GLM-5 · Zhipu AI's February 2026 open-weight frontier model: a 744B-parameter mixture-of-experts LLM released under the MIT license, followed by the GLM-5.2 update in June 2026.
- GPT-5 · OpenAI's August 2025 flagship: a unified system routing between fast and reasoning models, ending the GPT-4-era model picker.
- GPT-5.1 · OpenAI's November 2025 update to GPT-5, pairing a warmer default conversational style with adaptive reasoning in Instant and Thinking variants.
- GPT-5.2 · OpenAI's December 2025 update to the GPT-5 series, shipped in Instant, Thinking, and Pro variants amid intensified competition with Google's Gemini 3.
- GPT-5.6 · OpenAI's July 2026 frontier model series (Sol, Terra, Luna), first released as a restricted preview to US-government-vetted organizations in June 2026.
- gpt-oss · OpenAI's August 2025 open-weight models (120B and 20B MoE, Apache 2.0), its first open weights since GPT-2.
- Grok 3 · xAI's February 2025 flagship model, trained on the Colossus supercluster and released with a dedicated reasoning mode and the DeepSearch agent.
- Grok 4 · xAI's July 2025 flagship reasoning model, trained with large-scale reinforcement learning and released alongside a multi-agent Grok 4 Heavy variant.
- Grok 4.1 · xAI's November 2025 update to Grok 4, focused on emotional intelligence, creative writing, and reduced hallucinations, with a developer-reported first place on the LMArena text leaderboard at launch.
- Grok 4.5 · xAI's July 2026 frontier model, pitched as an efficient coding and agentic workhorse with configurable reasoning effort.
- Hunyuan 3.0 · Tencent's July 2026 open-weight flagship: a 295B-parameter mixture-of-experts model with 21B active parameters and a 256K context window, released under Apache 2.0.
- Kimi k1.5 · Moonshot AI's January 2025 multimodal reasoning model, trained with long-context reinforcement learning and released the same week as DeepSeek-R1.
- Kimi K3 · Moonshot AI's July 2026 flagship: a 2.8-trillion-parameter MoE, the largest open-weight release to date, with the Kimi Delta Attention architecture.
- Magistral · Mistral AI's first reasoning model line, released in June 2025 as the open-weight Magistral Small (24B) and the proprietary Magistral Medium.
- MiMo · Xiaomi's open model line: reasoning-first small models (MiMo-7B, 2025) scaled to the frontier checkpoints that led OpenRouter usage in 2026.
- MiniMax-M1 · MiniMax's June 2025 open-weight hybrid-attention reasoning model, developer-reported as the first large-scale open-weight model combining lightning attention with mixture-of-experts.
- MiniMax-M2 · MiniMax's October 2025 open-weight mixture-of-experts model aimed at agentic and coding workflows, released under the MIT license.
- o1 · OpenAI's first reasoning model (September 2024): trained with RL to think in long chains before answering, opening the test-time compute era.
- o3 · OpenAI's frontier reasoning model, previewed in December 2024 and released in April 2025, noted for developer-reported gains on ARC-AGI, math, and coding benchmarks.
- o4-mini · OpenAI's April 2025 small reasoning model, released alongside o3 with tool use and image-based reasoning at a lower price point.
- Qwen3 · Alibaba's April 2025 open-weight model family introducing hybrid thinking and non-thinking modes across dense and mixture-of-experts variants.
- QwQ-32B · Alibaba Qwen team's open-weight 32B reasoning model, previewed in November 2024 and released in March 2025 with developer-reported parity to much larger reasoners.
- Reinforcement learning with verifiable rewards · Post-training with programmatic reward signals (test suites, answer checkers) instead of learned reward models; the engine of the reasoning-model era.
- Test-time compute · Improving answers by spending more computation at inference (longer reasoning chains, search, sampling) rather than more pretraining; the axis behind reasoning models.
- Trinity Large Thinking · Arcee AI's April 2026 open-weight reasoning model: an Apache 2.0 variant of the 400B-parameter Trinity Large, post-trained with SFT and RL for long-horizon agent and tool-use workloads.
- Trinity Mini · Arcee AI's 26B-parameter sparse mixture-of-experts model (3B active) for agents and tool orchestration, released open-weight under Apache 2.0 in December 2025.