Category: Multimodal models
This category contains 43 pages.
- Amazon Nova · Amazon's in-house foundation-model family (December 2024): the Micro/Lite/Pro ladder plus creative models, distributed exclusively through Bedrock.
- Amazon Titan · Amazon's first-party foundation-model family (announced April 2023), served exclusively through Amazon Bedrock and later superseded by Amazon Nova.
- Apple foundation models · Apple's family of proprietary language models (an approximately 3B-parameter on-device model plus a larger server model) powering Apple Intelligence, announced June 2024.
- Aya · Cohere For AI's open multilingual model family (2024-2025), built with thousands of volunteer contributors to extend instruction-tuned LLMs beyond English.
- BLIP · Salesforce Research's vision-language pre-training family (2022-2023), whose BLIP-2 Q-Former popularized bridging frozen image encoders to frozen LLMs.
- Claude 3 · Anthropic's March 2024 generation (Haiku, Sonnet, Opus); Opus was the first model to consistently challenge GPT-4-class results.
- CLIP · OpenAI's January 2021 contrastive vision-language model that learns image classification from natural-language supervision and became a standard component in text-to-image systems.
- DALL-E · OpenAI's January 2021 text-to-image model, a 12-billion-parameter autoregressive Transformer that generated images from natural-language prompts.
- DALL-E 2 · OpenAI's April 2022 text-to-image diffusion model, which generated 1024x1024 images from CLIP embeddings and popularized AI image generation.
- DALL-E 3 · OpenAI's September 2023 text-to-image model, built to follow detailed prompts closely and integrated natively into ChatGPT.
- ERNIE 5 · Baidu's fifth-generation ERNIE flagship series: the natively omni-modal ERNIE 5.0 (November 2025) and the efficiency-focused ERNIE 5.1 (May 2026).
- Flamingo · DeepMind's April 2022 visual language model that pioneered few-shot multimodal prompting by bridging frozen vision and language models.
- FLUX · Black Forest Labs' family of rectified flow transformer text-to-image models, launched in August 2024 with open-weight and proprietary variants.
- Gemini 1.0 · Google DeepMind's December 2023 natively multimodal model family (Ultra, Pro, Nano), the first flagship release of the Gemini line.
- Gemini 1.5 · Google DeepMind's February 2024 generation; made million-token context windows a product reality and introduced the Flash tier.
- Gemini 2.0 · Google DeepMind's December 2024 'agentic era' generation: Flash-first release, native tool use, and the Flash Thinking reasoning experiments.
- Gemini 3.5 · Google DeepMind's 2026 Gemini generation: Flash shipped at I/O in May 2026 as the default Gemini model, while the Pro flagship remained unreleased as of July 2026.
- Genie · Google DeepMind's series of generative interactive world models (Genie, Genie 2, Genie 3) that produce playable environments from images or text.
- GPT-4 · OpenAI's fourth-generation GPT model, released March 14, 2023; multimodal input and undisclosed architecture.
- GPT-4 Turbo · OpenAI's November 2023 update to GPT-4 with a 128,000-token context window, fresher training data, and substantially lower API pricing.
- GPT-4o · OpenAI's May 2024 'omni' model: natively multimodal across text, vision, and audio, with real-time voice; became ChatGPT's long-running default.
- GPT-4o mini · OpenAI's July 2024 small multimodal model, a low-cost successor to GPT-3.5 Turbo in ChatGPT and the API.
- Hailuo · MiniMax's consumer AI brand, covering the Hailuo AI assistant app and the Hailuo video-generation model line (video-01, Hailuo 02).
- HunyuanVideo · Tencent's December 2024 open-weight text-to-video generation model, a roughly 13-billion-parameter diffusion transformer released with code and weights.
- IDEFICS · Hugging Face's August 2023 open reproduction of DeepMind's Flamingo: an 80B open-weight visual language model trained only on public data.
- iFlytek Spark · iFlytek's family of large language models, launched in May 2023 and notable for being trained on domestic Huawei Ascend compute.
- Imagen · Google's text-to-image diffusion model line, introduced in May 2022 and continued through Imagen 2, 3, and 4.
- InternVL · Shanghai AI Laboratory's open-weight vision-language model family, launched December 2023 and iterated through InternVL 1.5, 2, 2.5, and 3.
- Kimi k1.5 · Moonshot AI's January 2025 multimodal reasoning model, trained with long-context reinforcement learning and released the same week as DeepSeek-R1.
- Kling · Kuaishou's text-to-video generation model, launched in June 2024 as one of the first widely accessible answers to OpenAI's Sora.
- Llama 3.2 · Meta AI's September 2024 Llama release adding small on-device text models (1B, 3B) and the family's first vision-capable models (11B, 90B).
- Molmo · Ai2's September 2024 family of open-weight vision-language models trained on the human-annotated PixMo dataset rather than synthetic captions distilled from proprietary systems.
- Movie Gen · Meta AI's October 2024 suite of media foundation models for text-to-video and synchronized audio generation, shown as a research preview without a public release.
- PaliGemma · Google's open vision-language model of May 2024, pairing a SigLIP vision encoder with a Gemma language decoder and designed for fine-tuning on downstream tasks.
- Pixtral · Mistral AI's first natively multimodal model line, opened in September 2024 with the Apache-2.0 Pixtral 12B vision-language model.
- Qwen2.5-VL · Alibaba's January 2025 open-weight vision-language model series (3B to 72B) with document parsing, object grounding, long-video understanding, and GUI-agent capabilities.
- Seedream · ByteDance Seed's text-to-image generation model family (Seedream 2.0-4.0), noted for bilingual Chinese-English prompting and integrated image editing.
- SenseNova · SenseTime's foundation model family, launched in April 2023 and iterated through multimodal and reasoning-focused versions.
- Sora · OpenAI's text-to-video generation model, previewed in February 2024 and built on a diffusion transformer over spacetime patches.
- Sora 2 · OpenAI's September 2025 video-and-audio generation model, launched alongside a consumer Sora app with likeness-based 'cameos'.
- SynthID · Google DeepMind's watermarking system for labeling AI-generated images, audio, video, and text, first announced in August 2023.
- Veo · Google DeepMind's flagship text-to-video generation model series, first announced in May 2024 and extended with native audio in Veo 3.
- Vision Transformer · Transformer applied to images by treating fixed-size patches as tokens; introduced by Dosovitskiy et al. in 2020, now the standard vision encoder in multimodal LLMs.