ElevenLabs' most emotionally rich, expressive speech synthesis model. Natural, life-like speech with high emotional range and contextual understanding, and support for natural multi-speaker dialogue.
ElevenLabs models
Access 19 ElevenLabs models through the Oxyy unified API including Eleven v3, Dubbing v1, Dubbing v2. Compare pricing, context windows and capabilities between ElevenLabs models.
ElevenLabs tokens processed on Oxyy· daily, UTC
Models 19
Dubs video with automatic speaker detection across 29 languages.
End-to-end dubbing model that preserves voice and emotion across 92 languages.
English-only voice changer model.
English-only ultra-fast speech synthesis model.
ElevenLabs' fastest speech synthesis model, built for real-time applications and the Agents Platform. Balances speed and quality at half the price per character.
State-of-the-art multilingual voice changer. Converts an existing recording into another voice while preserving delivery.
State-of-the-art multilingual voice designer model.
Lifelike, consistent-quality speech synthesis. The most stable model for long-form generation, with consistent voice quality and accent across language switches.
ElevenLabs' speech models combined in one low-latency agent pipeline, for adding voice to a chat agent.
Royalty-free sound effects generation from text prompts.
Human-like and expressive voice DESIGN model. Generates a new voice from a text description rather than speech from text.
ElevenLabs' most expressive realtime speech synthesis model, tuned for low-latency conversation while keeping v3's emotional range.
Removes background noise, ambient sounds, reverb and interference to leave clean dialogue.
Studio-grade music generation from text prompts, composition plans and previously generated songs.
ElevenLabs' most advanced music model. Studio-grade generation from text prompts, composition plans and previously generated songs, with richer melodies, deeper arrangements and more layered instruments than music_v2.
State-of-the-art batch speech recognition across 90+ languages, with precise word-level timestamps, speaker diarization and dynamic audio tagging.
A fine-tune of Scribe v2 for medical and clinical audio. Improves recognition of drug names, anatomy, pathology and clinical dictation while matching Scribe v2 on everyday speech.
ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.
