Google: Gemini 3.8 Flash TTS

google/gemini-3.8-flash-tts

Google Gemini 3.8 Flash TTS API pricing

Gemini 3.8 Flash TTS costs $0.200 per million input tokens and $3.60 per million output tokens on Oxyy. These are the pay-as-you-go prices after the current 60% discount (Launch Promo; standard price $0.500 input and $9.00 output), and cached input tokens are billed at $0.050 per million. It takes text as input and returns audio. A cheaper Google option is Gemini 3.8 Flash-Lite TTS at $0.200 / $2.40 per 1M tokens; for every model and rate, see Gemini API pricing.

Google's flagship text-to-speech model for voice design, voice acting and long-form audio.
ModalitiesTextAudio
Price$0.200 / $3.60 per 1M tokens in / outAll Gemini API pricing
ReleasedSep 23, 2026

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
Google$0.200$3.60$0.0505.06s—

Capabilities (6)

StreamingVoice DesignVoice InstructionsMulti SpeakerPrompt cachingFree Tier

Supported parameters (4)

contentsspeechConfigresponseModalitiestemperature

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs on Oxyy, pay as you go. "Charged" is the rate a request is billed at; the standard price is shown beside it for reference.

Input price
$0.200
per 1M input tokens · standard price $0.500
Output price
$3.60
per 1M output tokens · standard price $9.00
RateStandard priceChargedUnit
Input$0.500$0.200per 1M input tokens
Output$9.00$3.60per 1M output tokens
Cached input$0.125$0.050per 1M cached input tokens
Cache storage$0.500$0.200per 1M tokens per hour stored

Discount. Charged rates include the current "Launch Promo" discount of 60% on every rate.

UTC

Specifications

Model ID
gemini-3.8-flash-tts
Developer
Google
Type
Text to speech
Input
Text
Output
Audio
Released
Sep 23, 2026
Tokenizer
Gemini
Series
Gemini 3.8
Endpoint
/v1/audio/speech

Best for

Text to speechVoice designMultilingual

Quickstart

Call gemini-3.8-flash-tts with a POST to /v1/audio/speech on https://api.oxyy.ai, using your Oxyy API key from the OXYY_API_KEY environment variable. The same key works for every model in the catalog.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.oxyy.ai/v1",
    api_key=os.environ["OXYY_API_KEY"],
)

with client.audio.speech.with_streaming_response.create(
    model="gemini-3.8-flash-tts",
    voice="alloy",
    input="Hello from Oxyy.",
    response_format="mp3",
) as response:
    response.stream_to_file("speech.mp3")

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
—tok/s
P50, streaming requests
Latency
5.06s
P50, end to end
UTC

Throughput

P99—
P95—
P50—

Latency

P504801ms
P758340ms
P9514.2s
P9919.9s

Time to first token

P50—
P95—

Cache hit rate

0.00%
Share of input tokens served from the prompt cache over the window.

Availability

The share of requests to this model that completed successfully over the last 30 days. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30 days)
56.60%
over 1,959 requests
Days with traffic
8
of the last 30 days

Days are UTC days. Days with no completed requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt134.8K
Completion601.1K

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is Gemini 3.8 Flash TTS?
Google's flagship text-to-speech model for voice design, voice acting and long-form audio.
How much does Gemini 3.8 Flash TTS cost?
Gemini 3.8 Flash TTS costs $0.200 per million input tokens and $3.60 per million output tokens on Oxyy. These are the pay-as-you-go prices after the current 60% discount (Launch Promo; standard price $0.500 input and $9.00 output), and cached input tokens are billed at $0.050 per million. You pay per request, with no subscription.
How do I call Gemini 3.8 Flash TTS through the API?
Send a POST request to https://api.oxyy.ai/v1/audio/speech with "model": "gemini-3.8-flash-tts", authenticated with your Oxyy API key. Put the text in "input" and choose a "voice"; the response body is the audio file.
What inputs and outputs does Gemini 3.8 Flash TTS support?
Gemini 3.8 Flash TTS accepts text as input and returns audio.
When was Gemini 3.8 Flash TTS released?
Gemini 3.8 Flash TTS was released on Sep 23, 2026.
What other models does Google have?
Google also offers Gemini 2.5 Flash TTS, Gemini 2.5 Pro TTS, Gemini 3.1 Flash TTS and Gemini 3.8 Flash-Lite TTS through Oxyy.

Alternatives to Gemini 3.8 Flash TTS

Models of the same kind from other vendors, closest in price first. Prices are pay-as-you-go on Oxyy, through the same API key.

Explore more models

More models from Google