Category: Mixture-of-experts models

This category contains 37 pages.

  • DBRX · Databricks' March 2024 open-weight mixture-of-experts language model, with 132 billion total and 36 billion active parameters.
  • DeepSeek-Coder-V2 · DeepSeek's June 2024 open-weight mixture-of-experts code model, developer-reported as competitive with closed frontier models on coding benchmarks.
  • DeepSeek-V2 · DeepSeek's May 2024 open MoE that introduced multi-head latent attention and DeepSeekMoE; started China's LLM price war.
  • DeepSeek-V3 · DeepSeek's December 2024 open-weight flagship: 671B-parameter MoE reporting frontier quality from a ~$5.6M disclosed final training run.
  • DeepSeek-V3.1 · DeepSeek's August 2025 hybrid open-weight model that merged thinking and non-thinking modes into a single checkpoint, succeeding both DeepSeek-V3 and DeepSeek-R1.
  • ERNIE 5 · Baidu's fifth-generation ERNIE flagship series: the natively omni-modal ERNIE 5.0 (November 2025) and the efficiency-focused ERNIE 5.1 (May 2026).
  • GLM-4.5 · Zhipu AI's July 2025 open-weight mixture-of-experts model, positioned as a unified agentic, reasoning, and coding foundation model.
  • GLM-4.6 · Zhipu AI's late-2025 open-weight mixture-of-experts model, an agentic-coding-focused successor to GLM-4.5 with a 200K-token context window.
  • GLM-5 · Zhipu AI's February 2026 open-weight frontier model: a 744B-parameter mixture-of-experts LLM released under the MIT license, followed by the GLM-5.2 update in June 2026.
  • gpt-oss · OpenAI's August 2025 open-weight models (120B and 20B MoE, Apache 2.0), its first open weights since GPT-2.
  • Grok-1 · xAI's first model (November 2023); its March 2024 open-weights drop, a 314B MoE under Apache 2.0, was then the largest open release.
  • Hunyuan · Tencent's model line: the open Hunyuan-Large MoE (2024), the leading open video model HunyuanVideo, and the Apache-2.0 Hunyuan 3.0 (2026).
  • Hunyuan 3.0 · Tencent's July 2026 open-weight flagship: a 295B-parameter mixture-of-experts model with 21B active parameters and a 256K context window, released under Apache 2.0.
  • Jamba · AI21 Labs' 2024 hybrid: the first production-scale model interleaving Transformer attention with Mamba state-space layers.
  • Kimi K2 · Moonshot AI's July 2025 trillion-parameter open-weight MoE; put open weights at the agentic frontier and introduced the MuonClip optimizer.
  • Kimi K3 · Moonshot AI's July 2026 flagship: a 2.8-trillion-parameter MoE, the largest open-weight release to date, with the Kimi Delta Attention architecture.
  • Ling (models) · Ant Group's open-weight mixture-of-experts model family (Bailing line), notable for frontier-scale training partly on Chinese domestic accelerators.
  • Llama 4 · Meta's April 2025 MoE generation (Scout, Maverick; Behemoth unreleased); a mixed reception that preceded Meta's Superintelligence Labs reorganization.
  • LongCat · Meituan's open MoE line: LongCat-Flash (560B, 2025) and LongCat-2.0 (1.6T, 2026), trained entirely on Chinese accelerators.
  • MAI-1 · Microsoft AI's in-house foundation model line: the mixture-of-experts MAI-1-preview entered public testing on LMArena in August 2025 and began feeding Copilot text features.
  • MiniMax-M1 · MiniMax's June 2025 open-weight hybrid-attention reasoning model, developer-reported as the first large-scale open-weight model combining lightning attention with mixture-of-experts.
  • MiniMax-M2 · MiniMax's October 2025 open-weight mixture-of-experts model aimed at agentic and coding workflows, released under the MIT license.
  • MiniMax-Text-01 · MiniMax's January 2025 open-weight language model, notable for its hybrid lightning-attention architecture and developer-reported 4-million-token context window.
  • Mixtral 8x22B · Mistral AI's April 2024 sparse mixture-of-experts model with about 141B total and 39B active parameters, released under Apache 2.0.
  • Mixtral 8x7B · Mistral AI's December 2023 sparse MoE, the first widely deployed open mixture-of-experts model; dropped as a magnet link before any announcement.
  • Qwen2 · Alibaba's June 2024 open-weight model series spanning 0.5B to 72B parameters, including one mixture-of-experts variant.
  • Qwen2.5-Max · Alibaba's January 2025 proprietary large-scale mixture-of-experts flagship, positioned against DeepSeek-V3 and Western frontier models.
  • Qwen3 · Alibaba's April 2025 open-weight model family introducing hybrid thinking and non-thinking modes across dense and mixture-of-experts variants.
  • Qwen3-Coder · Alibaba's July 2025 open-weight agentic coding model, a 480B-parameter mixture-of-experts specialist derived from the Qwen3 family.
  • Qwen3-Coder-Next · Alibaba Qwen team's February 2026 open-weight coding model, an ultra-sparse 80B hybrid-attention MoE tuned for coding agents and local development.
  • Qwen3-Max · Alibaba's September 2025 proprietary flagship of the Qwen3 generation, developer-reported to exceed one trillion parameters.
  • Step-2 · StepFun's 2024 trillion-parameter-scale mixture-of-experts language model, one of the first Chinese models announced at that scale.
  • Trinity (model family) · Arcee AI's open-weight sparse mixture-of-experts line: edge-scale Nano, agent-focused Mini, the 400B Trinity Large, and a Thinking reasoning branch.
  • Trinity Large · Arcee AI's 400B-parameter sparse mixture-of-experts flagship, trained from scratch on 17 trillion tokens and released in January 2026; widely described as the largest open-weight model from a US lab to that date.
  • Trinity Large Thinking · Arcee AI's April 2026 open-weight reasoning model: an Apache 2.0 variant of the 400B-parameter Trinity Large, post-trained with SFT and RL for long-horizon agent and tool-use workloads.
  • Trinity Mini · Arcee AI's 26B-parameter sparse mixture-of-experts model (3B active) for agents and tool orchestration, released open-weight under Apache 2.0 in December 2025.
  • Trinity Nano · Arcee AI's 6B-parameter sparse mixture-of-experts model (about 1B active per token) with a 128K context window, aimed at edge, embedded, and offline deployments.