Category: Models
This category contains 199 pages.
- AFM-4.5B · The first from-scratch Arcee Foundation Model: a 4.5B-parameter enterprise-focused open-weight LLM trained on 8 trillion tokens, later extended to a 64K context window.
- ALBERT · A 2019 parameter-efficient variant of BERT from Google Research that used factorized embeddings and cross-layer parameter sharing to cut model size.
- AlexNet · Krizhevsky, Sutskever, and Hinton's 2012 GPU-trained deep convolutional network whose runaway ILSVRC win (15.3 percent top-5 error) launched the deep learning era in computer vision.
- Alpaca · Stanford's March 2023 instruction-tuned LLaMA-7B, built for under $600 of synthetic data; proved frontier-adjacent assistants could be replicated for pocket change.
- AlphaFold · DeepMind's protein structure prediction line: CASP13 winner (2018), CASP14's effective solution of single-chain folding via AlphaFold 2 (2020), and AlphaFold 3 (2024) for biomolecular complexes.
- AlphaGo · DeepMind's Go system: deep policy and value networks guiding tree search, the first program to beat a Go professional (2015) and 4-1 winner over Lee Sedol in March 2016.
- AlphaZero · DeepMind's general self-play algorithm (2017) that mastered chess, shogi, and Go from the rules alone, defeating Stockfish, Elmo, and AlphaGo Zero without human data.
- Amazon Nova · Amazon's in-house foundation-model family (December 2024): the Micro/Lite/Pro ladder plus creative models, distributed exclusively through Bedrock.
- Amazon Titan · Amazon's first-party foundation-model family (announced April 2023), served exclusively through Amazon Bedrock and later superseded by Amazon Nova.
- Apple foundation models · Apple's family of proprietary language models (an approximately 3B-parameter on-device model plus a larger server model) powering Apple Intelligence, announced June 2024.
- Aya · Cohere For AI's open multilingual model family (2024-2025), built with thousands of volunteer contributors to extend instruction-tuned LLMs beyond English.
- BERT · Google's 2018 encoder-only Transformer, pretrained with masked language modeling; defined the pretrain-then-finetune era in NLP.
- BLIP · Salesforce Research's vision-language pre-training family (2022-2023), whose BLIP-2 Q-Former popularized bridging frozen image encoders to frozen LLMs.
- BLOOM · BigScience's 176-billion-parameter open-access multilingual language model, released in July 2022 as a large-scale collaborative science project coordinated by Hugging Face.
- Cerebras-GPT · Cerebras's March 2023 family of seven open compute-optimal GPT models (111M to 13B), trained on The Pile to demonstrate its wafer-scale hardware.
- ChatGLM · Zhipu AI and Tsinghua's 2023 bilingual open-weight chat assistant line, launched with the locally runnable ChatGLM-6B.
- Chinchilla · DeepMind's 2022 compute-optimal model: 70B parameters trained on 1.4T tokens, proving the GPT-3 generation was undertrained and resetting industry training budgets.
- Claude 1 · Anthropic's first publicly released chat assistant (March 2023), notable for its Constitutional AI alignment approach and later 100K-token context window.
- Claude 2 · Anthropic's July 2023 model; its 100K-token context window was the era's longest, and its cautious refusals defined early Claude's reputation.
- Claude 3 · Anthropic's March 2024 generation (Haiku, Sonnet, Opus); Opus was the first model to consistently challenge GPT-4-class results.
- Claude 3.5 Haiku · Anthropic's fast, low-cost model of the Claude 3.5 generation, announced in October 2024 as the successor to Claude 3 Haiku.
- Claude 3.5 Sonnet · Anthropic's mid-2024 workhorse; beat Claude 3 Opus at twice the speed, and its October refresh introduced computer use.
- Claude 3.7 Sonnet · Anthropic's February 2025 hybrid reasoning model: one network serving instant and extended-thinking modes, launched alongside Claude Code.
- Claude 4 · Anthropic's May 2025 generation (Opus 4, Sonnet 4), built for agentic coding; Opus 4 was the first model deployed under ASL-3 safeguards.
- Claude Fable 5 · Anthropic's June 2026 frontier model, first of the Mythos-class tier above Opus; briefly suspended under a US export-control directive.
- Claude Haiku 4.5 · Anthropic's October 2025 small model, the first Haiku with extended thinking, offering developer-reported near-Sonnet-4 coding performance at a fraction of the cost.
- Claude Instant · Anthropic's lighter, faster, lower-cost assistant line launched alongside the original Claude in March 2023 and retired after the Claude 3 generation introduced Haiku.
- Claude Opus 4.1 · Anthropic's August 2025 incremental upgrade to Claude Opus 4, with developer-reported gains in agentic coding and long-horizon reasoning.
- Claude Opus 4.5 · Anthropic's November 2025 flagship: developer-reported coding leadership with a major price cut, closing the Claude 4.x series.
- Claude Opus 4.6 · Anthropic's February 2026 flagship, adding a beta 1M-token context window, adaptive thinking, and effort controls to the Opus line.
- Claude Opus 4.8 · Anthropic's May 2026 Opus flagship, adding dynamic workflows, a discounted fast mode, and developer-reported gains in coding honesty; the last Opus release before Claude Fable 5.
- Claude Opus 5 · Anthropic's July 2026 Opus release, positioned as approaching Claude Fable 5's capability at half its price; the first Opus-tier model of the Claude 5 generation.
- Claude Sonnet 4.5 · Anthropic's September 2025 mid-tier hybrid reasoning model, marketed at launch as the company's strongest model for coding and computer use.
- Claude Sonnet 5 · Anthropic's June 2026 mid-tier model, marketed as its most agentic Sonnet, released as the default model for Free and Pro plans.
- CLIP · OpenAI's January 2021 contrastive vision-language model that learns image classification from natural-language supervision and became a standard component in text-to-image systems.
- Code Llama · Meta AI's August 2023 family of open-weight code-specialized language models derived from Llama 2, released in 7B to 70B sizes.
- Codestral · Mistral AI's May 2024 code-generation model, a 22-billion-parameter weight-available system trained on more than 80 programming languages.
- Codex (2021 model) · OpenAI's 2021 code-generation model, a GPT-3 descendant fine-tuned on public code; the original engine of GitHub Copilot and source of the HumanEval benchmark.
- Comma (models) · Pair of 7-billion-parameter open-weight language models released in June 2025 by EleutherAI and Common Pile collaborators, trained solely on the openly licensed Common Pile v0.1 corpus.
- Command (models) · Cohere's enterprise model line; Command R (2024) made retrieval-augmented generation and tool use the product, and Command A (2025) optimized for two-GPU serving.
- DALL-E · OpenAI's January 2021 text-to-image model, a 12-billion-parameter autoregressive Transformer that generated images from natural-language prompts.
- DALL-E 2 · OpenAI's April 2022 text-to-image diffusion model, which generated 1024x1024 images from CLIP embeddings and popularized AI image generation.
- DALL-E 3 · OpenAI's September 2023 text-to-image model, built to follow detailed prompts closely and integrated natively into ChatGPT.
- DBRX · Databricks' March 2024 open-weight mixture-of-experts language model, with 132 billion total and 36 billion active parameters.
- DeepSeek LLM · DeepSeek's first general-purpose open-weight language model series (7B and 67B), released in late 2023 and accompanied by a scaling-laws study.
- DeepSeek-Coder-V2 · DeepSeek's June 2024 open-weight mixture-of-experts code model, developer-reported as competitive with closed frontier models on coding benchmarks.
- DeepSeek-R1 · DeepSeek's January 2025 open-weight reasoning model; matched o1-class results with a published RL recipe and triggered a global market repricing.
- DeepSeek-V2 · DeepSeek's May 2024 open MoE that introduced multi-head latent attention and DeepSeekMoE; started China's LLM price war.
- DeepSeek-V3 · DeepSeek's December 2024 open-weight flagship: 671B-parameter MoE reporting frontier quality from a ~$5.6M disclosed final training run.
- DeepSeek-V3.1 · DeepSeek's August 2025 hybrid open-weight model that merged thinking and non-thinking modes into a single checkpoint, succeeding both DeepSeek-V3 and DeepSeek-R1.
- Devstral · Mistral AI's open-weight agentic coding model line, launched May 2025 with All Hands AI and tuned for software-engineering agent workflows.
- DistilBERT · Hugging Face's 2019 distilled version of BERT: roughly 40% smaller and 60% faster while retaining about 97% of BERT's language-understanding performance.
- Doubao · ByteDance's assistant model line; China's largest chatbot user base, priced to force market-wide cuts.
- ELECTRA · A 2020 Google Research and Stanford encoder model pre-trained with replaced token detection, reaching BERT-class quality at a fraction of the compute.
- ELMo · 2018 contextualized word embedding model from AI2 and the University of Washington, built on a bidirectional LSTM language model.
- ERNIE 5 · Baidu's fifth-generation ERNIE flagship series: the natively omni-modal ERNIE 5.0 (November 2025) and the efficiency-focused ERNIE 5.1 (May 2026).
- ERNIE Bot · Baidu's conversational AI service and flagship LLM line, launched in March 2023 as China's first major ChatGPT-style chatbot.
- Falcon (models) · The UAE Technology Innovation Institute's open model line; Falcon 180B was 2023's largest open release, and Falcon Mamba pioneered open SSM scale.
- Flamingo · DeepMind's April 2022 visual language model that pioneered few-shot multimodal prompting by bridging frozen vision and language models.
- Flan-T5 · Google's October 2022 family of instruction-finetuned T5 models, released openly in five sizes and a standard baseline for instruction-tuning research.
- FLUX · Black Forest Labs' family of rectified flow transformer text-to-image models, launched in August 2024 with open-weight and proprietary variants.
- Gemini 1.0 · Google DeepMind's December 2023 natively multimodal model family (Ultra, Pro, Nano), the first flagship release of the Gemini line.
- Gemini 1.5 · Google DeepMind's February 2024 generation; made million-token context windows a product reality and introduced the Flash tier.
- Gemini 2.0 · Google DeepMind's December 2024 'agentic era' generation: Flash-first release, native tool use, and the Flash Thinking reasoning experiments.
- Gemini 2.5 · Google DeepMind's 2025 thinking-by-default generation; 2.5 Pro led leaderboards through spring 2025.
- Gemini 3 · Google DeepMind's November 2025 generation; launch-led major benchmarks and deepened the Deep Think reasoning tier.
- Gemini 3.5 · Google DeepMind's 2026 Gemini generation: Flash shipped at I/O in May 2026 as the default Gemini model, while the Pro flagship remained unreleased as of July 2026.
- Gemma · Google DeepMind's open-weight line drawing on Gemini research: Gemma 1 (2024) through the multimodal Gemma 3 (2025).
- Genie · Google DeepMind's series of generative interactive world models (Genie, Genie 2, Genie 3) that produce playable environments from images or text.
- GLM-130B · A 130-billion-parameter open bilingual (English-Chinese) language model released by Tsinghua University and Zhipu AI in August 2022.
- GLM-4 · Zhipu AI's January 2024 flagship language model generation, spanning a proprietary API flagship and later open-weight GLM-4-9B variants.
- GLM-4.5 · Zhipu AI's July 2025 open-weight mixture-of-experts model, positioned as a unified agentic, reasoning, and coding foundation model.
- GLM-4.6 · Zhipu AI's late-2025 open-weight mixture-of-experts model, an agentic-coding-focused successor to GLM-4.5 with a 200K-token context window.
- GLM-5 · Zhipu AI's February 2026 open-weight frontier model: a 744B-parameter mixture-of-experts LLM released under the MIT license, followed by the GLM-5.2 update in June 2026.
- GloVe · Stanford's 2014 word-embedding method that fits vectors to global co-occurrence statistics, the main pre-Transformer alternative to word2vec.
- Gopher · DeepMind's December 2021 280-billion-parameter dense language model, the baseline that motivated the Chinchilla compute-optimal scaling analysis.
- GPT-1 · OpenAI's June 2018 Generative Pre-trained Transformer, the 117-million-parameter model that established the pretrain-then-finetune recipe behind the GPT series.
- GPT-2 · OpenAI's 1.5-billion-parameter 2019 model; demonstrated multitask behavior from pure language modeling and started the staged-release debate.
- GPT-3 · OpenAI's 175-billion-parameter language model of 2020, which demonstrated few-shot in-context learning at scale.
- GPT-3.5 · OpenAI's 2022 model series bridging GPT-3 and GPT-4; the original engine of ChatGPT.
- GPT-4 · OpenAI's fourth-generation GPT model, released March 14, 2023; multimodal input and undisclosed architecture.
- GPT-4 Turbo · OpenAI's November 2023 update to GPT-4 with a 128,000-token context window, fresher training data, and substantially lower API pricing.
- GPT-4.5 · OpenAI's February 2025 research-preview chat model, positioned as its largest scaling-focused (non-reasoning) model and retired from the API in mid-2025.
- GPT-4o · OpenAI's May 2024 'omni' model: natively multimodal across text, vision, and audio, with real-time voice; became ChatGPT's long-running default.
- GPT-4o mini · OpenAI's July 2024 small multimodal model, a low-cost successor to GPT-3.5 Turbo in ChatGPT and the API.
- GPT-5 · OpenAI's August 2025 flagship: a unified system routing between fast and reasoning models, ending the GPT-4-era model picker.
- GPT-5.1 · OpenAI's November 2025 update to GPT-5, pairing a warmer default conversational style with adaptive reasoning in Instant and Thinking variants.
- GPT-5.2 · OpenAI's December 2025 update to the GPT-5 series, shipped in Instant, Thinking, and Pro variants amid intensified competition with Google's Gemini 3.
- GPT-5.6 · OpenAI's July 2026 frontier model series (Sol, Terra, Luna), first released as a restricted preview to US-government-vetted organizations in June 2026.
- GPT-J · EleutherAI's June 2021 open-weight 6-billion-parameter autoregressive language model, an early open alternative to GPT-3-class systems.
- GPT-Neo · EleutherAI's March 2021 family of open-source autoregressive language models (up to 2.7 billion parameters), an early public replication of the GPT-3 design.
- GPT-NeoX-20B · EleutherAI's February 2022 open-source 20-billion-parameter autoregressive language model, at release the largest publicly available dense model with fully open weights.
- gpt-oss · OpenAI's August 2025 open-weight models (120B and 20B MoE, Apache 2.0), its first open weights since GPT-2.
- Grok 3 · xAI's February 2025 flagship model, trained on the Colossus supercluster and released with a dedicated reasoning mode and the DeepSearch agent.
- Grok 4 · xAI's July 2025 flagship reasoning model, trained with large-scale reinforcement learning and released alongside a multi-agent Grok 4 Heavy variant.
- Grok 4.1 · xAI's November 2025 update to Grok 4, focused on emotional intelligence, creative writing, and reduced hallucinations, with a developer-reported first place on the LMArena text leaderboard at launch.
- Grok 4.5 · xAI's July 2026 frontier model, pitched as an efficient coding and agentic workhorse with configurable reasoning effort.
- Grok-1 · xAI's first model (November 2023); its March 2024 open-weights drop, a 314B MoE under Apache 2.0, was then the largest open release.
- Grok-1.5 · xAI's March 2024 successor to Grok-1, adding a 128,000-token context window and developer-reported gains in math and coding.
- Grok-2 · xAI's August 2024 frontier chatbot model, released in beta on the X platform alongside a smaller Grok-2 mini variant.
- Hailuo · MiniMax's consumer AI brand, covering the Hailuo AI assistant app and the Hailuo video-generation model line (video-01, Hailuo 02).
- Hunyuan · Tencent's model line: the open Hunyuan-Large MoE (2024), the leading open video model HunyuanVideo, and the Apache-2.0 Hunyuan 3.0 (2026).
- Hunyuan 3.0 · Tencent's July 2026 open-weight flagship: a 295B-parameter mixture-of-experts model with 21B active parameters and a 256K context window, released under Apache 2.0.
- HunyuanVideo · Tencent's December 2024 open-weight text-to-video generation model, a roughly 13-billion-parameter diffusion transformer released with code and weights.
- IBM Granite · IBM's family of open-weight enterprise language models, spanning dense, mixture-of-experts, code, and hybrid Mamba/transformer variants.
- IDEFICS · Hugging Face's August 2023 open reproduction of DeepMind's Flamingo: an 80B open-weight visual language model trained only on public data.
- iFlytek Spark · iFlytek's family of large language models, launched in May 2023 and notable for being trained on domestic Huawei Ascend compute.
- Imagen · Google's text-to-image diffusion model line, introduced in May 2022 and continued through Imagen 2, 3, and 4.
- InstructGPT · OpenAI's 2022 RLHF-aligned GPT-3 variants; demonstrated that alignment training beats scale on human preference and set the template for ChatGPT.
- InternLM · Shanghai AI Laboratory's open-weight language model series, spanning the original 2023 release through InternLM2, InternLM2.5, and InternLM3.
- InternVL · Shanghai AI Laboratory's open-weight vision-language model family, launched December 2023 and iterated through InternVL 1.5, 2, 2.5, and 3.
- Jamba · AI21 Labs' 2024 hybrid: the first production-scale model interleaving Transformer attention with Mamba state-space layers.
- Jurassic (models) · AI21 Labs' family of large autoregressive language models, spanning Jurassic-1 (2021) and Jurassic-2 (2023), later succeeded by Jamba.
- Kimi k1.5 · Moonshot AI's January 2025 multimodal reasoning model, trained with long-context reinforcement learning and released the same week as DeepSeek-R1.
- Kimi K2 · Moonshot AI's July 2025 trillion-parameter open-weight MoE; put open weights at the agentic frontier and introduced the MuonClip optimizer.
- Kimi K3 · Moonshot AI's July 2026 flagship: a 2.8-trillion-parameter MoE, the largest open-weight release to date, with the Kimi Delta Attention architecture.
- Kling · Kuaishou's text-to-video generation model, launched in June 2024 as one of the first widely accessible answers to OpenAI's Sora.
- LaMDA · Google's 2021 dialogue-specialized language model family, a predecessor of Bard and the Gemini line, and the subject of a widely covered 2022 sentience controversy.
- Large language model · A neural language model, typically a decoder-only Transformer with billions of parameters, pretrained on internet-scale text.
- LFM (models) · Liquid AI's model line: the proprietary non-Transformer LFM1 generation (2024) and the open-weight LFM2/LFM2.5 edge families built on hybrid convolution-attention blocks.
- Ling (models) · Ant Group's open-weight mixture-of-experts model family (Bailing line), notable for frontier-scale training partly on Chinese domestic accelerators.
- LLaMA · Meta's February 2023 research models (7B-65B); their prompt leak seeded the open-source LLM ecosystem.
- Llama 2 · Meta's July 2023 openly licensed models (7B-70B) with RLHF chat variants; made open weights an official corporate strategy.
- Llama 3 · Meta's 2024 generation: 8B/70B in April, the 405B frontier-scale release in July; the herd paper documented open training at frontier scale.
- Llama 3.1 · Meta AI's July 2024 open-weight model family (8B, 70B, 405B), whose 405B variant was widely described as the first open-weight model competitive with frontier closed systems.
- Llama 3.2 · Meta AI's September 2024 Llama release adding small on-device text models (1B, 3B) and the family's first vision-capable models (11B, 90B).
- Llama 3.3 · Meta AI's December 2024 open-weight 70B instruct model, developer-reported to approach Llama 3.1 405B quality at a fraction of the serving cost.
- Llama 4 · Meta's April 2025 MoE generation (Scout, Maverick; Behemoth unreleased); a mixed reception that preceded Meta's Superintelligence Labs reorganization.
- LLaVA · The 2023 open vision-language recipe: a frozen vision encoder bridged to an open LLM with GPT-4-generated instruction data.
- LongCat · Meituan's open MoE line: LongCat-Flash (560B, 2025) and LongCat-2.0 (1.6T, 2026), trained entirely on Chinese accelerators.
- Magistral · Mistral AI's first reasoning model line, released in June 2025 as the open-weight Magistral Small (24B) and the proprietary Magistral Medium.
- MAI-1 · Microsoft AI's in-house foundation model line: the mixture-of-experts MAI-1-preview entered public testing on LMArena in August 2025 and began feeding Copilot text features.
- Meena · Google's January 2020 open-domain neural chatbot, a 2.6-billion-parameter conversational model that preceded LaMDA.
- Megatron-LM · NVIDIA's 2019 research project and open-source framework for training multi-billion-parameter Transformer language models with model parallelism.
- Megatron-Turing NLG · A 530-billion-parameter dense language model trained jointly by Microsoft and NVIDIA, announced in October 2021 as the largest dense Transformer of its time.
- MiMo · Xiaomi's open model line: reasoning-first small models (MiMo-7B, 2025) scaled to the frontier checkpoints that led OpenRouter usage in 2026.
- MiniMax-M1 · MiniMax's June 2025 open-weight hybrid-attention reasoning model, developer-reported as the first large-scale open-weight model combining lightning attention with mixture-of-experts.
- MiniMax-M2 · MiniMax's October 2025 open-weight mixture-of-experts model aimed at agentic and coding workflows, released under the MIT license.
- MiniMax-Text-01 · MiniMax's January 2025 open-weight language model, notable for its hybrid lightning-attention architecture and developer-reported 4-million-token context window.
- Mistral 7B · Mistral AI's September 2023 debut model: a 7.3-billion-parameter open-weight language model that outperformed larger Llama 2 variants.
- Mistral Large · Mistral AI's flagship proprietary language model line, launched February 2024 and updated in July 2024 as the 123-billion-parameter Mistral Large 2.
- Mistral Medium 3 · Mistral AI's May 2025 proprietary enterprise model, marketed as near-frontier performance at a substantially lower price.
- Mistral Small · Mistral AI's mid-size model line, launched February 2024 and later released open-weight, positioned for low-latency deployment below Mistral Large.
- Mixtral 8x22B · Mistral AI's April 2024 sparse mixture-of-experts model with about 141B total and 39B active parameters, released under Apache 2.0.
- Mixtral 8x7B · Mistral AI's December 2023 sparse MoE, the first widely deployed open mixture-of-experts model; dropped as a magnet link before any announcement.
- Molmo · Ai2's September 2024 family of open-weight vision-language models trained on the human-annotated PixMo dataset rather than synthetic captions distilled from proprietary systems.
- Movie Gen · Meta AI's October 2024 suite of media foundation models for text-to-video and synchronized audio generation, shown as a research preview without a public release.
- MPT (models) · MosaicML's 2023 open-weight MPT line (7B and 30B), an early commercially usable alternative to research-only LLaMA licenses.
- Nemotron · NVIDIA's family of open-weight language models, best known for the June 2024 Nemotron-4 340B release aimed at synthetic data generation.
- NETtalk · Sejnowski and Rosenberg's 1986 network that learned to pronounce English text by backpropagation, the parallel distributed processing era's most public demonstration that multi-layer networks could learn a realistic cognitive task.
- o1 · OpenAI's first reasoning model (September 2024): trained with RL to think in long chains before answering, opening the test-time compute era.
- o3 · OpenAI's frontier reasoning model, previewed in December 2024 and released in April 2025, noted for developer-reported gains on ARC-AGI, math, and coding benchmarks.
- o4-mini · OpenAI's April 2025 small reasoning model, released alongside o3 with tool use and image-based reasoning at a lower price point.
- OLMo · AI2's fully open model line: weights, data, code, and checkpoints all public; the scientific control group of the LLM era.
- OPT · Meta AI's May 2022 suite of open decoder-only language models, from 125M to 175B parameters, released with weights, code, and a candid training logbook.
- PaliGemma · Google's open vision-language model of May 2024, pairing a SigLIP vision encoder with a Gemma language decoder and designed for fine-tuning on downstream tasks.
- PaLM · Google's 540-billion-parameter Pathways Language Model of April 2022, a landmark in dense scaling and few-shot reasoning.
- PaLM 2 · Google's May 2023 large language model, successor to PaLM, that powered Bard and Workspace features before being replaced by Gemini.
- Pixtral · Mistral AI's first natively multimodal model line, opened in September 2024 with the Apache-2.0 Pixtral 12B vision-language model.
- Pythia · EleutherAI's April 2023 suite of 16 open-weight language models (70M-12B parameters) with released training checkpoints, built for research on training dynamics and interpretability.
- Qwen 1 · Alibaba Cloud's first generation of Qwen large language models (2023), released as open-weight checkpoints from 1.8B to 72B parameters.
- Qwen2 · Alibaba's June 2024 open-weight model series spanning 0.5B to 72B parameters, including one mixture-of-experts variant.
- Qwen2.5 · Alibaba's September 2024 open-weight model series (0.5B to 72B), trained on a developer-reported 18 trillion tokens and widely used as a base for fine-tunes and reasoning distillations.
- Qwen2.5-Max · Alibaba's January 2025 proprietary large-scale mixture-of-experts flagship, positioned against DeepSeek-V3 and Western frontier models.
- Qwen2.5-VL · Alibaba's January 2025 open-weight vision-language model series (3B to 72B) with document parsing, object grounding, long-video understanding, and GUI-agent capabilities.
- Qwen3 · Alibaba's April 2025 open-weight model family introducing hybrid thinking and non-thinking modes across dense and mixture-of-experts variants.
- Qwen3-Coder · Alibaba's July 2025 open-weight agentic coding model, a 480B-parameter mixture-of-experts specialist derived from the Qwen3 family.
- Qwen3-Coder-Next · Alibaba Qwen team's February 2026 open-weight coding model, an ultra-sparse 80B hybrid-attention MoE tuned for coding agents and local development.
- Qwen3-Max · Alibaba's September 2025 proprietary flagship of the Qwen3 generation, developer-reported to exceed one trillion parameters.
- QwQ-32B · Alibaba Qwen team's open-weight 32B reasoning model, previewed in November 2024 and released in March 2025 with developer-reported parity to much larger reasoners.
- RoBERTa · Facebook AI's July 2019 replication study of BERT that showed the original model was significantly undertrained, setting new benchmark records with an optimized pretraining recipe.
- RWKV · Community-built linear-attention RNN trained like a Transformer; the open ecosystem's longest-running alternative-architecture project.
- Seed-OSS · ByteDance Seed's August 2025 open-weight 36B dense language model family, notable for a long native context window and a user-controllable thinking budget.
- Seedream · ByteDance Seed's text-to-image generation model family (Seedream 2.0-4.0), noted for bilingual Chinese-English prompting and integrated image editing.
- Segment Anything · Meta AI's April 2023 promptable image segmentation model (SAM), released with the 1.1-billion-mask SA-1B dataset.
- SenseNova · SenseTime's foundation model family, launched in April 2023 and iterated through multimodal and reasoning-focused versions.
- SmolLM · Hugging Face's family of small open-weight language models (135M to 3B parameters), launched in July 2024 and built on openly documented training corpora.
- Sora · OpenAI's text-to-video generation model, previewed in February 2024 and built on a diffusion transformer over spacetime patches.
- Sora 2 · OpenAI's September 2025 video-and-audio generation model, launched alongside a consumer Sora app with likeness-based 'cameos'.
- Stable Diffusion · Open-weight latent diffusion text-to-image model family released in August 2022 by Stability AI with CompVis and Runway, spanning v1 through Stable Diffusion 3.5.
- StarCoder · BigCode's open code models (2023-2024), trained on the consent-audited Stack corpus; the open-governance benchmark for code LLMs.
- Step-2 · StepFun's 2024 trillion-parameter-scale mixture-of-experts language model, one of the first Chinese models announced at that scale.
- SynthID · Google DeepMind's watermarking system for labeling AI-generated images, audio, video, and text, first announced in August 2023.
- T5 · Google's 2019 Text-to-Text Transfer Transformer, an open-weight encoder-decoder family that cast every NLP task as text-to-text.
- TD-Gammon · Gerald Tesauro's IBM backgammon program that taught itself master-level play by temporal-difference self-play, the first landmark success of reinforcement learning on a complex task.
- Trinity Large · Arcee AI's 400B-parameter sparse mixture-of-experts flagship, trained from scratch on 17 trillion tokens and released in January 2026; widely described as the largest open-weight model from a US lab to that date.
- Trinity Large Thinking · Arcee AI's April 2026 open-weight reasoning model: an Apache 2.0 variant of the 400B-parameter Trinity Large, post-trained with SFT and RL for long-horizon agent and tool-use workloads.
- Trinity Mini · Arcee AI's 26B-parameter sparse mixture-of-experts model (3B active) for agents and tool orchestration, released open-weight under Apache 2.0 in December 2025.
- Trinity Nano · Arcee AI's 6B-parameter sparse mixture-of-experts model (about 1B active per token) with a 128K context window, aimed at edge, embedded, and offline deployments.
- Tulu 3 · Ai2's November 2024 family of openly post-trained Llama 3.1 derivatives, released with its full recipe and known for introducing reinforcement learning with verifiable rewards (RLVR).
- Turing-NLG · Microsoft's 17-billion-parameter language model of February 2020, briefly the largest published language model and the public debut of the DeepSpeed library.
- ULMFiT · Howard and Ruder's 2018 three-stage transfer-learning recipe for text, a direct precursor to the pretrain-and-fine-tune paradigm of BERT and GPT.
- Veo · Google DeepMind's flagship text-to-video generation model series, first announced in May 2024 and extended with native audio in Veo 3.
- Vicuna · March 2023 open chatbot from the LMSYS team, fine-tuned from LLaMA on shared ChatGPT conversations and a catalyst for LLM-as-a-judge evaluation.
- Whisper · OpenAI's open-source speech recognition model of September 2022, trained on 680,000 hours of weakly supervised multilingual audio.
- XLNet · A June 2019 generalized autoregressive pretraining model from Google Brain and Carnegie Mellon University that outperformed BERT on many benchmarks.
- YaLM-100B · Yandex's June 2022 open 100-billion-parameter GPT-like language model, at release the largest dense model with freely downloadable Apache 2.0 weights.
- YandexGPT · Family of large language models by Yandex, launched May 2023 inside the Alice assistant and served via the Yandex Cloud API, with openly released 8B checkpoints.
- Yi (model family) · 01.AI's family of bilingual open-weight language models, launched in November 2023 with Yi-6B and Yi-34B and later extended by Yi-1.5 and the proprietary Yi-Large.