Main Page
Backpropagation is the algorithm that computes the gradient of a neural network's loss with respect to every parameter, applying the chain rule backward through the computation graph in a single pass that costs roughly as much as the forward pass itself. Paired with a gradient-based optimizer, it is the training procedure of essentially every modern network, from the perceptron revival of the 1980s to today's Transformers. Popularized by Rumelhart, Hinton, and Williams in 1986, the method was developed by Paul Werbos in the early 1970s, with antecedents in the adjoint equations of 1960s optimal control theory. (Full article...)
| 2017 | The Transformer is introduced in "Attention Is All You Need". |
|---|---|
| 2018 | BERT establishes pretraining plus fine-tuning as the dominant recipe. |
| 2019 | GPT-2 scales decoder-only generation and starts the staged-release debate. |
| 2020 | GPT-3 demonstrates few-shot in-context learning; scaling laws formalize the compute-performance frontier. |
| 2022 | InstructGPT brings RLHF to production models; ChatGPT takes language models mainstream. |
| 2023 | GPT-4 sets a new capability bar; LLaMA ignites the open-weight ecosystem. |
| 2024 | GPT-4o goes natively multimodal; o1 shifts the frontier toward test-time compute. |
| 2025 | DeepSeek-R1 triggers the DeepSeek R1 release shock; Claude 4 and GPT-5 contest the frontier. |
| 2026 | Claude Fable 5 and GPT-5.6 ship under brief export-control restrictions; open weights reach trillion-parameter scale with Kimi K3 and LongCat, trained increasingly on Chinese silicon. |
- ... that Dropout, part of the original Transformer recipe, quietly left frontier pretraining once single-epoch training on trillion-token corpora removed the overfitting it targets?
- ... that The Pile, the 825 GiB corpus that trained GPT-J and Pythia, drew its conversational text from sources including Ubuntu IRC logs and the Enron email corpus?
- ... that roughly 60 percent of GPT-3's weighted training tokens came from a filtered snapshot of Common Crawl, a nonprofit's free archive of the public web?
- ... that Safe Superintelligence raised a reported $2 billion at a valuation of roughly $32 billion in 2025, while it had yet to release a product?
- ... that Cerebras Systems fabricates its Wafer-Scale Engine as a single silicon wafer, the first generation carrying 400,000 cores and 1.2 trillion transistors?
- ... that Sam Altman learned of his removal from OpenAI on a Google Meet call, with five to ten minutes' notice, while attending the Las Vegas Grand Prix?
- ... that attention heads park surplus attention weight on a sequence's first tokens, and that keeping roughly four of these attention sinks cached lets models generate stably over millions of tokens?
Large language model · Transformer (architecture) · Self-attention · Feed-forward network · Layer normalization · RMSNorm · Mixture of experts · Scaling laws · Tokenization · Next-token prediction · Pretraining · In-context learning · Supervised fine-tuning · Instruction tuning · Reward model · Reinforcement learning from human feedback · Direct preference optimization · Reinforcement learning with verifiable rewards · Constitutional AI · Chain-of-thought prompting · Test-time compute · Distributed training
- All pages: every article on one page
- Categories: browse by topic
- Architecture: interactive timeline of model design
- July 24: Anthropic releases Claude Opus 5, described by the developer as approaching Claude Fable 5's capability at half its price; it becomes the default model on the Claude Max plan.
- July 17: Moonshot AI announces Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, described by the developer as the largest open-weight release to date.
- July: Meituan releases LongCat-2.0, a 1.6-trillion-parameter open model reported as trained entirely on Chinese-made accelerators.
- July 9: OpenAI begins the broad rollout of GPT-5.6 (variants Sol, Terra, and Luna) after a restricted preview limited to government-vetted organizations.
- June 30: Anthropic releases Claude Sonnet 5 and makes it the default model for its Free and Pro plans.
- June 9: Anthropic announces Claude Fable 5, the first of its Mythos-class tier; a US export-control directive suspends access on June 12 and is lifted at the end of the month.
- February 11: Zhipu AI releases GLM-5, a 744-billion-parameter mixture-of-experts model, as open weights under the MIT license.
2025–2026: GPT-5 · Claude Opus 4.5 · Claude 4 · DeepSeek-R1 · Claude 3.7 Sonnet
2023–2024: GPT-4 · Llama 2 · LLaMA · DeepSeek-V3 · GPT-4o · o1 · Claude 3.5 Sonnet · Claude 3 · Llama 3
2018–2022: BERT · GPT-2 · GPT-3 · InstructGPT
OpenAI · Anthropic · Google DeepMind · Meta AI · Ndea · xAI · Thinking Machines Lab · Cursor · Safe Superintelligence · Prometheus · Mistral AI · Cohere · NVIDIA · Microsoft AI · Yandex · Hugging Face · EleutherAI · Ai2 · Arcee AI · Stability AI · Databricks · Cerebras Systems · AI21 Labs · Aleph Alpha · Technology Innovation Institute · Nous Research · Liquid AI · Moondream
Research laboratories: DeepSeek · Moonshot AI · Zhipu AI · MiniMax · StepFun · Baichuan · 01.AI
Technology companies: Alibaba (Qwen) · ByteDance · Tencent · Baidu · Huawei · Xiaomi · Ant Group · Meituan
Research institutes: Shanghai AI Laboratory · iFlytek · SenseTime