swap_horiz Voicebox Alternatives
Looking for alternatives to Voicebox? Compare the top Model options ranked by our AI scoring system.
Voicebox
Voicebox is a generative artificial intelligence model for speech synthesis developed by Meta and announced in 2023. Utilizing a non-autoregressive flow-matching architecture, it is capable of zero-shot text-to-speech generation in multiple languages, including English, French, Spanish, German, Poli...
apps Top Voicebox Alternatives
The top alternative to Voicebox in 2026 is WaveNet with a score of 9.1/10, followed by SAM (9.1) and Llama 3.1 405B (8.9).
WaveNet
WaveNet is a deep neural network architecture for generating raw audio waveforms, introduced by researchers at DeepMind...
SAM
SAM (Segment Anything Model) is a vision foundation model developed by Meta AI and released in 2023. It was trained on t...
Llama 3.1 405B
Llama 3.1 405B is a large language model released by Meta in 2024, serving as the flagship of the Llama 3.1 collection....
Tacotron 2
Tacotron 2 is a neural network architecture for text-to-speech (TTS) synthesis introduced by Google researchers in 2017....
DINOv2
DINOv2 is a self-supervised vision foundation model developed by Meta AI and released in 2023. It was trained on a highl...
SAM 2
SAM 2 (Segment Anything Model 2) is an artificial intelligence model developed by Meta and released in 2024. It extends...
Llama 3.3
Llama 3.3 is an instruction-tuned large language model developed by Meta and released in December 2024. Built with 70 bi...
Embed v3
Embed v3 is a generation of text embedding models developed by the enterprise artificial intelligence company Cohere, re...
ElevenLabs Turbo v2.5
ElevenLabs Turbo v2.5 is a low-latency multilingual text-to-speech model from ElevenLabs (2024), optimized for real-time...
Mamba
Mamba is a deep learning architecture introduced in 2023 by researchers Albert Gu and Tri Dao that utilizes selective st...
AlphaCode 2
AlphaCode 2 is an artificial intelligence system developed by Google DeepMind for competitive programming, announced in...
Suno v3
Suno v3 is an artificial intelligence music generation model developed by Suno and released in December 2023. The model...
Imagen 2
Imagen 2 is a text-to-image diffusion model developed by Google DeepMind, released in late 2023 as the successor to the...
Mistral 7B
Mistral 7B is an open-weight large language model released in September 2023 by the French artificial intelligence compa...
SigLIP
SigLIP (Sigmoid Loss for Language Image Pre-training) is a vision-language model introduced by Google in 2023. It modifi...
MusicGen
MusicGen is a text-to-music generation model developed by Meta and released in 2023 as part of the AudioCraft open-sourc...
Gemini Ultra
Gemini Ultra is the largest model in Google DeepMind's first-generation Gemini family, introduced alongside Gemini Pro a...
AudioLM
AudioLM is a framework for audio generation introduced by Google Research in 2022. It models both speech and music by fi...
SoundStorm
SoundStorm is a neural network model developed by Google DeepMind and detailed in 2023 that specializes in high-quality...
Emu Video
Emu Video is a text-to-video generation model introduced by Meta in 2023. It utilizes a factorized, or cascaded, diffusi...
summarize Quick Comparison Summary
See all Model ranked by score
emoji_events View Full Model Rankings