Best Vision Language
No tags available
Rankings use category fit, feature coverage, pricing signals, public reception, and recency. Affiliate relationships do not affect scores.
CLIP (Contrastive Language–Image Pretraining) is a neural network model introduced by OpenAI in 2021. It is trained on approximately four hundred million image and text pairs collected from the internet using a contrastive objective that aligns image and text representations in a shared embedding sp...
Qwen2-VL is a vision-language model developed by Alibaba as part of the Qwen series, released in 2024. The model is designed to process visual and textual data, featuring a Naive Dynamic Resolution mechanism that allows it to natively handle images and videos of varying sizes without forced cropping...
InternVL2 is an open-source vision-language foundation model developed by the Shanghai AI Laboratory, released in 2024. It is designed to process and reason across both visual and textual data, integrating a vision encoder with a large language model. The architecture is available in various paramet...
Flamingo is a multimodal visual language model introduced by DeepMind in 2022. The architecture is designed to process arbitrarily interleaved sequences of images and text, allowing it to perform visual question answering and image captioning. It achieves strong few-shot learning capabilities by con...
SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision-language model introduced by Google in 2023. It modifies standard contrastive learning frameworks by replacing the typical softmax loss with a pairwise sigmoid loss. This architectural change removes the need for global comparisons ac...
ALIGN (Large-scale ImaGe and Noisy-text embedding) is a vision-language model developed by Google Research in 2021. The model utilizes a dual-encoder architecture and is trained using contrastive learning on a massive dataset of over one billion noisy image-text pairs collected from the web without...
LLaVA 1.6 is an open-weight, vision-language model developed by researchers from the University of Wisconsin–Madison and collaborating institutions. Released in early 2024, this iteration improves upon LLaVA 1.5 by supporting higher-resolution image inputs, which significantly enhances its optical c...
CogVLM2 is an open-source, multimodal vision-language model developed through a collaboration between Zhipu AI and Tsinghua University. Released in 2024, the architecture is designed for high-resolution image processing, specifically supporting inputs up to 1344 x 1344 pixels. It serves as a researc...
LLaVA 1.5 is an open-weight multimodal large language model developed by researchers at the University of Wisconsin-Madison and released in 2023. The architecture connects a CLIP vision encoder to the Vicuna language model through an MLP projection layer, enabling the model to process and reason abo...
GLM-4V is a multimodal vision-language model developed by Zhipu AI in 2024 as part of the GLM-4 model family. The architecture extends the GLM-4 language model with image understanding capabilities, allowing it to process visual information and answer questions about images in both Chinese and Engli...
You're in. We'll email you when new Vision Language entries land.