Best 2023 Model
No tags available
Rankings use category fit, feature coverage, pricing signals, public reception, and recency. Affiliate relationships do not affect scores.
SAM (Segment Anything Model) is a vision foundation model developed by Meta AI and released in 2023. It was trained on the SA-1B dataset of over one billion segmentation masks on eleven million images, enabling zero-shot generalization to segment objects not present in training data. Users can promp...
DINOv2 is a self-supervised vision foundation model developed by Meta AI and released in 2023. It was trained on a highly curated dataset of 142 million images without relying on manual labels or text supervision. By utilizing an improved student-teacher architecture, the model produces robust visua...
Mamba is a deep learning architecture introduced in 2023 by researchers Albert Gu and Tri Dao that utilizes selective state space models (SSMs) for natural language processing. Unlike traditional Transformer models that require quadratic computational complexity for sequence length, Mamba achieves l...
Embed v3 is a generation of text embedding models developed by the enterprise artificial intelligence company Cohere, released in late 2023. The models are specifically designed to enhance retrieval-augmented generation (RAG) systems by mapping text into dense vector representations for semantic sea...
AlphaCode 2 is an artificial intelligence system developed by Google DeepMind for competitive programming, announced in 2023. It utilizes the Gemini large language model architecture and employs a reinforcement learning pipeline to generate, evaluate, and filter potential code solutions. During test...
Mistral 7B is an open-weight large language model released in September 2023 by the French artificial intelligence company Mistral AI. The model features 7 billion parameters and utilizes architectural efficiencies like Grouped-Query Attention and Sliding Window Attention to accelerate processing an...
SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision-language model introduced by Google in 2023. It modifies standard contrastive learning frameworks by replacing the typical softmax loss with a pairwise sigmoid loss. This architectural change removes the need for global comparisons ac...
Voicebox is a generative artificial intelligence model for speech synthesis developed by Meta and announced in 2023. Utilizing a non-autoregressive flow-matching architecture, it is capable of zero-shot text-to-speech generation in multiple languages, including English, French, Spanish, German, Poli...
Imagen 2 is a text-to-image diffusion model developed by Google DeepMind, released in late 2023 as the successor to the original Imagen model. It was engineered to deliver enhanced photorealism, improved image-text alignment, and significantly better rendering of text within generated images, such a...
MusicGen is a text-to-music generation model developed by Meta and released in 2023 as part of the AudioCraft open-source framework. Built on an EnCodec tokenizer and a transformer-based language model architecture, it generates audio waveforms from text prompts. The model is capable of producing mu...
Gemini Ultra is the largest model in Google DeepMind's first-generation Gemini family, introduced alongside Gemini Pro and Gemini Nano in December 2023 with general availability following in early 2024. It is a multimodal model designed to process text, images, audio, and video within a single archi...
Vicuna-13B is an open-weight large language model released in 2023 by LMSYS Org, a research organization comprising members from UC Berkeley, CMU, Stanford, and UCSD. The model was created by fine-tuning Meta's LLaMA-13B architecture on approximately 70,000 user-shared conversations sourced from Sha...
SoundStorm is a neural network model developed by Google DeepMind and detailed in 2023 that specializes in high-quality audio generation. It operates by predicting audio token sequences in a parallel, non-autoregressive manner, a design that allows it to synthesize audio significantly faster than re...
ViT-22B is a vision transformer model developed by Google Research and released in 2023. With 22 billion parameters, it represented one of the largest vision transformer architectures at the time of publication, demonstrating how scaling laws that had been established for language models might also...
Runway Gen-2 is an AI model released by the company Runway in 2023, designed for text-to-video and image-to-video generation. The model allows users to create short video clips from text prompts or by animating existing images, using a diffusion-based approach to generate temporally consistent conte...
LLaVA 1.5 is an open-weight multimodal large language model developed by researchers at the University of Wisconsin-Madison and released in 2023. The architecture connects a CLIP vision encoder to the Vicuna language model through an MLP projection layer, enabling the model to process and reason abo...
Pythia is a suite of 16 open-source language models developed by EleutherAI in 2023, ranging from 70 million to 12 billion parameters. The models were trained on identical data sequences with publicly available checkpoints at regular intervals, specifically designed to enable scientific study of lar...
Yi-34B is a 34-billion-parameter language model developed by 01.AI, an artificial intelligence company founded by Kai-Fu Lee, and released in 2023. The model is explicitly designed as bilingual, with strong capabilities in both English and Chinese across reasoning, mathematics, coding, and general k...
OpenHermes 2.5 is a 2023 instruction-tuned language model released by Nous Research and derived from the seven-billion-parameter Mistral 7B model. It was trained on a large mixture of primarily synthetic instruction and conversation data, including responses produced by GPT-4. The model is intended...
Alpaca is an instruction-following language model released by Stanford researchers in March 2023. It was fine-tuned from Meta's LLaMA 7B model using 52,000 instruction-response examples generated with OpenAI's text-davinci-003 model through a self-instruct process. The project demonstrated that a re...
Qwen-VL is a vision-language model introduced by Alibaba Cloud as a multimodal extension of the Qwen language-model family. It can process images together with text for tasks including image captioning, visual question answering, text recognition, and dialogue about visual content. A distinguishing...
Solar 10.7B is an open-weight large language model developed by the South Korean artificial intelligence company Upstage and released in 2023. It was constructed using a technique called depth up-scaling, which involves extending the architecture of Meta's Llama 2 model to create a deeper, more capa...
ERNIE 4.0 is a large language model and artificial intelligence system developed by the Chinese technology company Baidu, officially unveiled in 2023. The acronym stands for Enhanced Representation through Knowledge Integration, reflecting the model's architecture, which is designed to integrate vas...
PaLM 2 is a large language model developed by Google, officially announced in 2023 as the successor to the original PaLM. It is designed with enhanced capabilities in multilingual understanding, reasoning, and coding. This model serves as the underlying technology for several Google products, includ...
MiniGPT-4 is a vision-language model introduced in 2023 by researchers at King Abdullah University of Science and Technology (KAUST). It aligns a frozen visual encoder derived from BLIP-2 with a frozen Vicuna large language model using a single linear projection layer trained on a comparatively smal...
Nous Hermes 2 is a series of instruction-tuned large language models released in 2024 by the independent research collective Nous Research. The models are fine-tuned from open-weight base models including Mixtral 8x7B and Meta's Llama family, using curated synthetic and human-generated datasets desi...
Emu Video is a text-to-video generation model introduced by Meta in 2023. It utilizes a factorized, or cascaded, diffusion-based approach that breaks the generation process into two distinct steps: creating an image based on a text prompt, and subsequently generating a video conditioned on both the...
Phi-2 is a 2.7-billion-parameter transformer language model developed by Microsoft, released in December 2023. It was trained using a methodology emphasizing curated high-quality data, including synthetic data and filtered web content, to improve reasoning capabilities relative to model size. Phi-2...
VideoPoet is a large language model developed by Google Research and presented in 2023 as a system for zero-shot video generation. The model is designed to handle video, image, audio, and text within a single unified architecture, enabling capabilities such as text-to-video generation, image-to-vide...
Falcon 180B is a 180-billion-parameter causal decoder-only large language model developed by the Technology Innovation Institute in the United Arab Emirates. Released in 2023 under a license that permits open-access commercial use, it was trained on 3.5 trillion tokens primarily drawn from the Refin...
You're in. We'll email you when new 2023 Model entries land.