swap_horiz LLaVA 1.5 Alternatives
Looking for alternatives to LLaVA 1.5? Compare the top Model options ranked by our AI scoring system.
LLaVA 1.5
LLaVA 1.5 is an open-weight multimodal large language model developed by researchers at the University of Wisconsin-Madison and released in 2023. The architecture connects a CLIP vision encoder to the Vicuna language model through an MLP projection layer, enabling the model to process and reason abo...
apps Top LLaVA 1.5 Alternatives
The top alternative to LLaVA 1.5 in 2026 is Qwen2-VL with a score of 8.6/10, followed by Flamingo (8.5) and InternVL2 (8.5).
Qwen2-VL
Qwen2-VL is a vision-language model developed by Alibaba as part of the Qwen series, released in 2024. The model is desi...
Flamingo
Flamingo is a multimodal visual language model introduced by DeepMind in 2022. The architecture is designed to process a...
InternVL2
InternVL2 is an open-source vision-language foundation model developed by the Shanghai AI Laboratory, released in 2024....
Mistral 7B
Mistral 7B is an open-weight large language model released in September 2023 by the French artificial intelligence compa...
SigLIP
SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision-language model introduced by Google in 2023. It modifi...
Gemini Ultra
Gemini Ultra is the largest model in Google DeepMind's first-generation Gemini family, introduced alongside Gemini Pro a...
LLaVA 1.6
LLaVA 1.6 is an open-weight, vision-language model developed by researchers from the University of Wisconsin–Madison and...
Llama 3.2
Llama 3.2 is a family of open-weight artificial intelligence models released by Meta in 2024. The release introduces lig...
Yi-34B
Yi-34B is a 34-billion-parameter language model developed by 01.AI, an artificial intelligence company founded by Kai-Fu...
GLM-4V
GLM-4V is a multimodal vision-language model developed by Zhipu AI in 2024 as part of the GLM-4 model family. The archit...
Vicuna 13B
Vicuna-13B is an open-weight large language model released in 2023 by LMSYS Org, a research organization comprising memb...
CogVLM2
CogVLM2 is an open-source, multimodal vision-language model developed through a collaboration between Zhipu AI and Tsing...
Qwen-VL
Qwen-VL is a vision-language model introduced by Alibaba Cloud as a multimodal extension of the Qwen language-model fami...
Idefics2
Idefics2 is an open-weight vision-language model family released by Hugging Face in 2024. It accepts combinations of ima...
Emu Video
Emu Video is a text-to-video generation model introduced by Meta in 2023. It utilizes a factorized, or cascaded, diffusi...
ERNIE 4.0
ERNIE 4.0 is a large language model and artificial intelligence system developed by the Chinese technology company Baidu...
Solar 10.7B
Solar 10.7B is an open-weight large language model developed by the South Korean artificial intelligence company Upstage...
Nous Hermes 2
Nous Hermes 2 is a series of instruction-tuned large language models released in 2024 by the independent research collec...
MiniGPT-4
MiniGPT-4 is a vision-language model introduced in 2023 by researchers at King Abdullah University of Science and Techno...
Kosmos-2
Microsoft's 2023 multimodal model that grounds natural-language phrases to specific image regions, enabling referring ex...
summarize Quick Comparison Summary
See all Model ranked by score
emoji_events View Full Model Rankings