The pool of Main Page "Did you know" hooks: 27 entries. Seven rotate onto the Main Page each build, keyed to the build date. Every hook restates a fact sourced in its featured article.

  • ... that attention heads park surplus attention weight on a sequence's first tokens, and that keeping roughly four of these attention sinks cached lets models generate stably over millions of tokens?
  • ... that Chinchilla matched Gopher's compute budget with a quarter of the parameters and four times the data, and outperformed it across the evaluation suite?
  • ... that the January 2025 DeepSeek R1 release shock erased roughly $593 billion of NVIDIA's market value in a single trading day?
  • ... that BLOOM was trained on a French public supercomputer, on a corpus spanning 46 natural languages and 13 programming languages?
  • ... that GLM-130B could run in INT4 precision on four consumer RTX 3090 GPUs, with little degradation reported by its developers?
  • ... that the Comma models, trained solely on the openly licensed Common Pile corpus, were reported to perform comparably to same-size models trained on unlicensed text?
  • ... that FlashAttention computes exact attention faster by minimizing GPU memory traffic rather than floating-point operations?
  • ... that OpenAI reportedly skipped the name "o2" and called its o1 successor o3 to avoid conflict with the O2 telecommunications trademark?
  • ... that Claude Fable 5 was suspended days after its June 2026 launch under the first direct US export-control action against deployed frontier models rather than hardware?
  • ... that Moonshot AI reported zero loss spikes across the training run of its roughly one-trillion-parameter model Kimi K2, a disclosure that made its MuonClip optimizer an immediate research topic?
  • ... that the Maverick variant of Meta AI's Llama 4 submitted to the Chatbot Arena leaderboard was an unreleased conversational tune, drawing a public rebuke from the leaderboard's maintainers?
  • ... that MiniMax reported that the full reinforcement-learning run for its reasoning model MiniMax-M1 used 512 Nvidia H800 GPUs for about three weeks, at a rental cost of roughly US$534,700?
  • ... that Claude 1, the first publicly released AI assistant from Anthropic, was announced on the same day OpenAI announced GPT-4?
  • ... that EleutherAI trained all 16 Pythia models on exactly the same data in exactly the same order, and released 154 intermediate training checkpoints for each?
  • ... that Meta AI released a day-by-day training logbook for its GPT-3-scale model OPT, documenting hardware failures, loss divergences, and mid-run restarts?
  • ... that Yandex's YaLM-100B, trained on a corpus that was 75% Russian-language, was at release the largest dense language model with weights downloadable under a permissive open-source license?
  • ... that Vicuna-13B, fine-tuned on shared ChatGPT conversations for a stated cost of around $300, was developer-reported to reach about 90% of ChatGPT's quality when judged by GPT-4?
  • ... that the earliest documented uses of step-by-step elicitation behind chain-of-thought prompting came not from a lab but from players of the GPT-3-backed game AI Dungeon in July 2020?
  • ... that clamping a single sparse autoencoder feature turned Claude 3 Sonnet into "Golden Gate Claude", a briefly public variant that steered conversations toward the bridge?
  • ... that PagedAttention, modeled on 1960s-era operating-system paging, cut wasted key-value cache memory in language model serving from as much as 80 percent to under 4 percent?
  • ... that teacher forcing, coined in 1989 for recurrent networks, remains the way every language model is pretrained?
  • ... that Dropout, part of the original Transformer recipe, quietly left frontier pretraining once single-epoch training on trillion-token corpora removed the overfitting it targets?
  • ... that The Pile, the 825 GiB corpus that trained GPT-J and Pythia, drew its conversational text from sources including Ubuntu IRC logs and the Enron email corpus?
  • ... that roughly 60 percent of GPT-3's weighted training tokens came from a filtered snapshot of Common Crawl, a nonprofit's free archive of the public web?
  • ... that Safe Superintelligence raised a reported $2 billion at a valuation of roughly $32 billion in 2025, while it had yet to release a product?
  • ... that Cerebras Systems fabricates its Wafer-Scale Engine as a single silicon wafer, the first generation carrying 400,000 cores and 1.2 trillion transistors?
  • ... that Sam Altman learned of his removal from OpenAI on a Google Meet call, with five to ten minutes' notice, while attending the Las Vegas Grand Prix?