oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. ElevenLabs/
  4. Scribe v2 Realtime
E

ElevenLabs: Scribe v2 Realtime

elevenlabs/scribe_v2_realtime
Compare Playground Get API key
ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.
ModalitiesAudioText
In / out price— / — per 1M
Context—
Released—
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
ElevenLabs—————

Capabilities (6)

StreamingRealtimeWord TimestampsLanguage DetectionEntity DetectionVoice Activity Detection

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
—
/M tokens
Effective output price
—
/M tokens
RateListedChargedUnit
UTC
Price history starts building from the first daily snapshot — there is not enough of it yet to draw a trend.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
—tok/s
P50, streaming requests
Latency
—s
P50, end to end
UTC
No requests to this model in the selected window, so there is nothing to measure yet.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
0.00%
over 1 requests
Days with traffic
1
of the last 30

Days are UTC days. Days with no requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

No apps have identified themselves for this model yet. Clients can opt in by sending an X-Title header with their request.

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt0
Completion0

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is Scribe v2 Realtime?
ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.
How much does Scribe v2 Realtime cost?
Pricing for Scribe v2 Realtime is listed on this page under Pricing. You pay per request, with no subscription.
What is the context length of Scribe v2 Realtime?
The context window for Scribe v2 Realtime is not published.
Does Scribe v2 Realtime support tool calling and structured outputs?
Neither tool calling nor structured outputs are listed among this model's capabilities. Any unsupported parameter you send is ignored rather than rejected.
What inputs and outputs does Scribe v2 Realtime support?
Scribe v2 Realtime accepts audio as input and returns text.
What other models does ElevenLabs have?
ElevenLabs also offers Dubbing v1, Dubbing v2, Eleven English v2 (Voice Changer), Eleven Flash v2 through Oxyy.
When was Scribe v2 Realtime released?
No release date is published for Scribe v2 Realtime.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from ElevenLabs

Dubbing v1Dubs video with automatic speaker detection across 29 languages.text to video · — contextDubbing v2End-to-end dubbing model that preserves voice and emotion across 92 languages.text to video · — contextEleven English v2 (Voice Changer)English-only voice changer model.audio tts · — contextEleven Flash v2English-only ultra-fast speech synthesis model.audio tts · — contextEleven Flash v2.5ElevenLabs' fastest speech synthesis model, built for real-time applications and the Agents Platform. Balances speed an…audio tts · — contextEleven Multilingual v2 (Voice Changer)State-of-the-art multilingual voice changer. Converts an existing recording into another voice while preserving deliver…audio tts · — context
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout