oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. ElevenLabs
E

ElevenLabs models

Access 19 ElevenLabs models through the Oxyy unified API including Eleven v3, Dubbing v1, Dubbing v2. Compare pricing, context windows and capabilities between ElevenLabs models.

ElevenLabs tokens processed on Oxyy· daily, UTC

Models 19

EElevenLabs: Eleven v3Text→Audio-60%7mo ago
text-to-speechexpressiveaudiobookmulti-speakeraudio-tags

ElevenLabs' most emotionally rich, expressive speech synthesis model. Natural, life-like speech with high emotional range and contextual understanding, and support for natural multi-speaker dialogue.

by ElevenLabsFeb 1, 2026— context—/M input—/M output5085ms latency
EElevenLabs: Dubbing v1AudioVideo→AudioVideo-60%—
dubbingvideospeaker-detection

Dubs video with automatic speaker detection across 29 languages.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Dubbing v2AudioVideo→AudioVideo-60%—
dubbingvideovoice-preservation

End-to-end dubbing model that preserves voice and emotion across 92 languages.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven English v2 (Voice Changer)Audio→Audio-60%—
voice-changerspeech-to-speechenglish-only

English-only voice changer model.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Flash v2Text→Audio-60%—
text-to-speechlow-latencyenglish-only

English-only ultra-fast speech synthesis model.

by ElevenLabs— context—/M input—/M output513ms latency
EElevenLabs: Eleven Flash v2.5Text→Audio-60%—
text-to-speechlow-latencyrealtimebulkvoice-agents

ElevenLabs' fastest speech synthesis model, built for real-time applications and the Agents Platform. Balances speed and quality at half the price per character.

by ElevenLabs— context—/M input—/M output430ms latency
EElevenLabs: Eleven Multilingual v2 (Voice Changer)Audio→Audio-60%—
voice-changerspeech-to-speechmultilingual

State-of-the-art multilingual voice changer. Converts an existing recording into another voice while preserving delivery.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Multilingual v2 (Voice Design)Text→Audio-60%—
voice-designtext-to-voicemultilingual

State-of-the-art multilingual voice designer model.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Multilingual v2Text→Audio-60%—
text-to-speechlong-formstablenarration

Lifelike, consistent-quality speech synthesis. The most stable model for long-form generation, with consistent voice quality and accent across language switches.

by ElevenLabs— context—/M input—/M output1961ms latency
EElevenLabs: ElevenAgents Speech EngineTextAudio→Audio-60%—
voice-agentsrealtimespeech-to-speechpipeline

ElevenLabs' speech models combined in one low-latency agent pipeline, for adding voice to a chat agent.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Sound Effects v2Text→Audio-60%—
sound-effectsaudio-generationroyalty-free

Royalty-free sound effects generation from text prompts.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven v3 (Voice Design / Text to Voice)Text→Audio-60%—
voice-designtext-to-voice

Human-like and expressive voice DESIGN model. Generates a new voice from a text description rather than speech from text.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven v3 ConversationalText→Audio-60%—
text-to-speechrealtimeexpressivevoice-agentsaudio-tags

ElevenLabs' most expressive realtime speech synthesis model, tuned for low-latency conversation while keeping v3's emotional range.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Voice IsolatorAudio→Audio-60%—
audio-processingnoise-removal

Removes background noise, ambient sounds, reverb and interference to leave clean dialogue.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Music v2TextAudio→Audio-60%—
music-generationtext-to-music

Studio-grade music generation from text prompts, composition plans and previously generated songs.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Music v2.5TextAudio→Audio-60%—
music-generationtext-to-musicstudio-grade

ElevenLabs' most advanced music model. Studio-grade generation from text prompts, composition plans and previously generated songs, with richer melodies, deeper arrangements and more layered instruments than music_v2.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Scribe v2Audio→Text-60%—
speech-to-textdiarizationtimestampsentity-detection

State-of-the-art batch speech recognition across 90+ languages, with precise word-level timestamps, speaker diarization and dynamic audio tagging.

by ElevenLabs— context—/M input—/M output1055ms latency
EElevenLabs: Scribe v2 MedicalAudio→Text-60%—
speech-to-textmedicalclinicalhipaadiarization

A fine-tune of Scribe v2 for medical and clinical audio. Improves recognition of drug names, anatomy, pathology and clinical dictation while matching Scribe v2 on everyday speech.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Scribe v2 RealtimeAudio→Text-60%—
speech-to-textrealtimestreaminglow-latencyvoice-agents

ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.

by ElevenLabs— context—/M input—/M output
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout