Category: Open-source AI

This category contains 31 pages.

  • Ai2 · Seattle nonprofit AI research institute (AI2); publisher of the fully open OLMo models, Dolma corpus, and Tulu post-training recipes.
  • BLIP · Salesforce Research's vision-language pre-training family (2022-2023), whose BLIP-2 Q-Former popularized bridging frozen image encoders to frozen LLMs.
  • BLOOM · BigScience's 176-billion-parameter open-access multilingual language model, released in July 2022 as a large-scale collaborative science project coordinated by Hugging Face.
  • Cohere · Toronto-based enterprise AI lab founded in 2019 by Transformer co-author Aidan Gomez; builds the Command model line, the North agent platform, and the Aya open-science program.
  • Common Crawl · A nonprofit's free archive of web crawls, petabytes of raw page data refreshed monthly, that serves as the raw substrate for most large language model training corpora.
  • Common Pile · EleutherAI's June 2025 corpus of 8 TB of public-domain and openly licensed text from 30 sources, released with the Comma v0.1 models as an openly licensed successor to the Pile.
  • Databricks · American data and AI platform company founded by the creators of Apache Spark, acquirer of MosaicML and publisher of the open-weight DBRX model.
  • EleutherAI · Grassroots research collective turned nonprofit; produced GPT-J, GPT-NeoX, The Pile, and the Pythia interpretability suite.
  • GloVe · Stanford's 2014 word-embedding method that fits vectors to global co-occurrence statistics, the main pre-Transformer alternative to word2vec.
  • GPT-J · EleutherAI's June 2021 open-weight 6-billion-parameter autoregressive language model, an early open alternative to GPT-3-class systems.
  • GPT-Neo · EleutherAI's March 2021 family of open-source autoregressive language models (up to 2.7 billion parameters), an early public replication of the GPT-3 design.
  • GPT-NeoX-20B · EleutherAI's February 2022 open-source 20-billion-parameter autoregressive language model, at release the largest publicly available dense model with fully open weights.
  • Hugging Face · The hub of the open machine-learning ecosystem: model hosting, the transformers library, datasets, and leaderboards.
  • Keras · Open-source deep learning API created by François Chollet in March 2015, rewritten in 2023 to run on JAX, TensorFlow, and PyTorch, and used as launch tooling for Gemma and other open LLMs.
  • Liquid AI · MIT CSAIL spinoff founded in 2023 by the liquid-neural-network researchers; builds the efficiency-first LFM model line for edge and on-device use, backed by a $250M AMD-led round.
  • Megatron-LM · NVIDIA's 2019 research project and open-source framework for training multi-billion-parameter Transformer language models with model parallelism.
  • Moondream · M87 Labs' tiny vision-language model line; the reference small VLM for edge and high-volume deployment.
  • MosaicML · San Francisco startup (2021-2023) whose efficient-training platform and open-weight MPT models led to a $1.3 billion Databricks acquisition; its team became Databricks Mosaic Research and built DBRX.
  • Nous Research · US open-source AI collective turned company; builds the Hermes instruction-tuned model line, the YaRN context-extension method, and the DisTrO/Psyche decentralized-training stack.
  • Pythia · EleutherAI's April 2023 suite of 16 open-weight language models (70M-12B parameters) with released training checkpoints, built for research on training dynamics and interpretability.
  • RoBERTa · Facebook AI's July 2019 replication study of BERT that showed the original model was significantly undertrained, setting new benchmark records with an optimized pretraining recipe.
  • SmolLM · Hugging Face's family of small open-weight language models (135M to 3B parameters), launched in July 2024 and built on openly documented training corpora.
  • Stability AI · Company behind the Stable Diffusion image models; emblem of the 2022 open generative-AI wave and its business-model turbulence.
  • Stable Diffusion · Open-weight latent diffusion text-to-image model family released in August 2022 by Stability AI with CompVis and Runway, spanning v1 through Stable Diffusion 3.5.
  • Technology Innovation Institute · Abu Dhabi state applied-research institute under the ATRC; pretrained the open Falcon model line on RefinedWeb, including Falcon 180B.
  • The Pile · EleutherAI's December 2020 corpus of 825 GiB of curated English text drawn from 22 sources, the training set behind GPT-Neo, GPT-J, GPT-NeoX-20B, and Pythia, later partly withdrawn amid the Books3 copyright dispute.
  • Tulu 3 · Ai2's November 2024 family of openly post-trained Llama 3.1 derivatives, released with its full recipe and known for introducing reinforcement learning with verifiable rewards (RLVR).
  • Vicuna · March 2023 open chatbot from the LMSYS team, fine-tuned from LLaMA on shared ChatGPT conversations and a catalyst for LLM-as-a-judge evaluation.
  • Whisper · OpenAI's open-source speech recognition model of September 2022, trained on 680,000 hours of weakly supervised multilingual audio.
  • YaLM-100B · Yandex's June 2022 open 100-billion-parameter GPT-like language model, at release the largest dense model with freely downloadable Apache 2.0 weights.
  • Yandex · Russian internet company behind the dominant Russian search engine and Alice assistant; released YaLM-100B openly in 2022 and builds the YandexGPT line.