ModelRefs / Best Audio AI Models in 2026
Best Audio AI Models in 2026
Speech-to-text, text-to-speech and music generation models. Whisper, ElevenLabs and open-source alternatives compared.
Overview
Audio models cover transcription (Whisper, Deepgram), TTS (ElevenLabs, OpenAI TTS) and music generation (MusicGen, Suno).
ModelRefs tracks 6 audio models with a published reference page. Each page states the provider, the capabilities ModelRefs has evidence for, the benchmark results it holds with their dates and sources, and the limitations and coverage gaps that remain.
Models in this category
6 audio models have a published reference page on ModelRefs.
Listed alphabetically. Category membership means a model is a candidate worth evaluating for this kind of work, not a ranking and not an endorsement — compare the evidence on each page against your own requirements.
Other model categories
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Audio AI Models in 2026.
Frequently asked questions
What is the best speech-to-text model?
OpenAI Whisper (large-v3) is the leading open-source ASR model. For low latency at scale, Deepgram and AssemblyAI APIs are popular.