Category: 2023 model releases

This category contains 38 pages.

  • Alpaca · Stanford's March 2023 instruction-tuned LLaMA-7B, built for under $600 of synthetic data; proved frontier-adjacent assistants could be replicated for pocket change.
  • Amazon Titan · Amazon's first-party foundation-model family (announced April 2023), served exclusively through Amazon Bedrock and later superseded by Amazon Nova.
  • BLIP · Salesforce Research's vision-language pre-training family (2022-2023), whose BLIP-2 Q-Former popularized bridging frozen image encoders to frozen LLMs.
  • Cerebras-GPT · Cerebras's March 2023 family of seven open compute-optimal GPT models (111M to 13B), trained on The Pile to demonstrate its wafer-scale hardware.
  • ChatGLM · Zhipu AI and Tsinghua's 2023 bilingual open-weight chat assistant line, launched with the locally runnable ChatGLM-6B.
  • Claude 1 · Anthropic's first publicly released chat assistant (March 2023), notable for its Constitutional AI alignment approach and later 100K-token context window.
  • Claude 2 · Anthropic's July 2023 model; its 100K-token context window was the era's longest, and its cautious refusals defined early Claude's reputation.
  • Claude Instant · Anthropic's lighter, faster, lower-cost assistant line launched alongside the original Claude in March 2023 and retired after the Claude 3 generation introduced Haiku.
  • Code Llama · Meta AI's August 2023 family of open-weight code-specialized language models derived from Llama 2, released in 7B to 70B sizes.
  • DALL-E 3 · OpenAI's September 2023 text-to-image model, built to follow detailed prompts closely and integrated natively into ChatGPT.
  • DeepSeek LLM · DeepSeek's first general-purpose open-weight language model series (7B and 67B), released in late 2023 and accompanied by a scaling-laws study.
  • ERNIE Bot · Baidu's conversational AI service and flagship LLM line, launched in March 2023 as China's first major ChatGPT-style chatbot.
  • Falcon (models) · The UAE Technology Innovation Institute's open model line; Falcon 180B was 2023's largest open release, and Falcon Mamba pioneered open SSM scale.
  • Gemini 1.0 · Google DeepMind's December 2023 natively multimodal model family (Ultra, Pro, Nano), the first flagship release of the Gemini line.
  • GPT-4 · OpenAI's fourth-generation GPT model, released March 14, 2023; multimodal input and undisclosed architecture.
  • GPT-4 Turbo · OpenAI's November 2023 update to GPT-4 with a 128,000-token context window, fresher training data, and substantially lower API pricing.
  • Grok-1 · xAI's first model (November 2023); its March 2024 open-weights drop, a 314B MoE under Apache 2.0, was then the largest open release.
  • IDEFICS · Hugging Face's August 2023 open reproduction of DeepMind's Flamingo: an 80B open-weight visual language model trained only on public data.
  • iFlytek Spark · iFlytek's family of large language models, launched in May 2023 and notable for being trained on domestic Huawei Ascend compute.
  • InternLM · Shanghai AI Laboratory's open-weight language model series, spanning the original 2023 release through InternLM2, InternLM2.5, and InternLM3.
  • InternVL · Shanghai AI Laboratory's open-weight vision-language model family, launched December 2023 and iterated through InternVL 1.5, 2, 2.5, and 3.
  • Jurassic (models) · AI21 Labs' family of large autoregressive language models, spanning Jurassic-1 (2021) and Jurassic-2 (2023), later succeeded by Jamba.
  • LLaMA · Meta's February 2023 research models (7B-65B); their prompt leak seeded the open-source LLM ecosystem.
  • Llama 2 · Meta's July 2023 openly licensed models (7B-70B) with RLHF chat variants; made open weights an official corporate strategy.
  • LLaVA · The 2023 open vision-language recipe: a frozen vision encoder bridged to an open LLM with GPT-4-generated instruction data.
  • Mistral 7B · Mistral AI's September 2023 debut model: a 7.3-billion-parameter open-weight language model that outperformed larger Llama 2 variants.
  • Mixtral 8x7B · Mistral AI's December 2023 sparse MoE, the first widely deployed open mixture-of-experts model; dropped as a magnet link before any announcement.
  • MPT (models) · MosaicML's 2023 open-weight MPT line (7B and 30B), an early commercially usable alternative to research-only LLaMA licenses.
  • PaLM 2 · Google's May 2023 large language model, successor to PaLM, that powered Bard and Workspace features before being replaced by Gemini.
  • Pythia · EleutherAI's April 2023 suite of 16 open-weight language models (70M-12B parameters) with released training checkpoints, built for research on training dynamics and interpretability.
  • Qwen 1 · Alibaba Cloud's first generation of Qwen large language models (2023), released as open-weight checkpoints from 1.8B to 72B parameters.
  • Segment Anything · Meta AI's April 2023 promptable image segmentation model (SAM), released with the 1.1-billion-mask SA-1B dataset.
  • SenseNova · SenseTime's foundation model family, launched in April 2023 and iterated through multimodal and reasoning-focused versions.
  • StarCoder · BigCode's open code models (2023-2024), trained on the consent-audited Stack corpus; the open-governance benchmark for code LLMs.
  • SynthID · Google DeepMind's watermarking system for labeling AI-generated images, audio, video, and text, first announced in August 2023.
  • Vicuna · March 2023 open chatbot from the LMSYS team, fine-tuned from LLaMA on shared ChatGPT conversations and a catalyst for LLM-as-a-judge evaluation.
  • YandexGPT · Family of large language models by Yandex, launched May 2023 inside the Alice assistant and served via the Yandex Cloud API, with openly released 8B checkpoints.
  • Yi (model family) · 01.AI's family of bilingual open-weight language models, launched in November 2023 with Yi-6B and Yi-34B and later extended by Yi-1.5 and the proprietary Yi-Large.