Category: Vision-language models
This category contains 2 pages.
- LLaVA · The 2023 open vision-language recipe: a frozen vision encoder bridged to an open LLM with GPT-4-generated instruction data.
- Moondream (model family) · The tiny-VLM line from M87 Labs: Moondream 1 (2024) through the MoE Moondream 3 Preview, built for edge and high-volume vision work.