Main Page

From LLM Wiki, the language model encyclopedia
Welcome to LLM Wiki, the language model encyclopedia. 361 articles · built July 24, 2026.
History at a glance
2017The Transformer is introduced in "Attention Is All You Need".
2018BERT establishes pretraining plus fine-tuning as the dominant recipe.
2019GPT-2 scales decoder-only generation and starts the staged-release debate.
2020GPT-3 demonstrates few-shot in-context learning; scaling laws formalize the compute-performance frontier.
2022InstructGPT brings RLHF to production models; ChatGPT takes language models mainstream.
2023GPT-4 sets a new capability bar; LLaMA ignites the open-weight ecosystem.
2024GPT-4o goes natively multimodal; o1 shifts the frontier toward test-time compute.
2025DeepSeek-R1 triggers the DeepSeek R1 release shock; Claude 4 and GPT-5 contest the frontier.
2026Claude Fable 5 and GPT-5.6 ship under brief export-control restrictions; open weights reach trillion-parameter scale with Kimi K3 and LongCat, trained increasingly on Chinese silicon.
Did you know ...
  • ... that Dropout, part of the original Transformer recipe, quietly left frontier pretraining once single-epoch training on trillion-token corpora removed the overfitting it targets?
  • ... that The Pile, the 825 GiB corpus that trained GPT-J and Pythia, drew its conversational text from sources including Ubuntu IRC logs and the Enron email corpus?
  • ... that roughly 60 percent of GPT-3's weighted training tokens came from a filtered snapshot of Common Crawl, a nonprofit's free archive of the public web?
  • ... that Safe Superintelligence raised a reported $2 billion at a valuation of roughly $32 billion in 2025, while it had yet to release a product?
  • ... that Cerebras Systems fabricates its Wafer-Scale Engine as a single silicon wafer, the first generation carrying 400,000 cores and 1.2 trillion transistors?
  • ... that Sam Altman learned of his removal from OpenAI on a Google Meet call, with five to ten minutes' notice, while attending the Las Vegas Grand Prix?
  • ... that attention heads park surplus attention weight on a sequence's first tokens, and that keeping roughly four of these attention sinks cached lets models generate stably over millions of tokens?

Show another set · Full hook archive

Getting started
This year in machine learning
  • July 24: Anthropic releases Claude Opus 5, described by the developer as approaching Claude Fable 5's capability at half its price; it becomes the default model on the Claude Max plan.
  • July 17: Moonshot AI announces Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, described by the developer as the largest open-weight release to date.
  • July: Meituan releases LongCat-2.0, a 1.6-trillion-parameter open model reported as trained entirely on Chinese-made accelerators.
  • July 9: OpenAI begins the broad rollout of GPT-5.6 (variants Sol, Terra, and Luna) after a restricted preview limited to government-vetted organizations.
  • June 30: Anthropic releases Claude Sonnet 5 and makes it the default model for its Free and Pro plans.
  • June 9: Anthropic announces Claude Fable 5, the first of its Mythos-class tier; a US export-control directive suspends access on June 12 and is lifted at the end of the month.
  • February 11: Zhipu AI releases GLM-5, a 744-billion-parameter mixture-of-experts model, as open weights under the MIT license.
Major models

2025–2026: GPT-5 · Claude Opus 4.5 · Claude 4 · DeepSeek-R1 · Claude 3.7 Sonnet

2023–2024: GPT-4 · Llama 2 · LLaMA · DeepSeek-V3 · GPT-4o · o1 · Claude 3.5 Sonnet · Claude 3 · Llama 3

2018–2022: BERT · GPT-2 · GPT-3 · InstructGPT

AI organizations in China
Main article: Chinese AI labs

Research laboratories: DeepSeek · Moonshot AI · Zhipu AI · MiniMax · StepFun · Baichuan · 01.AI

Technology companies: Alibaba (Qwen) · ByteDance · Tencent · Baidu · Huawei · Xiaomi · Ant Group · Meituan

Research institutes: Shanghai AI Laboratory · iFlytek · SenseTime