Audio API
Text to speech
ElevenLabs, OpenAI TTS and Gemini TTS, through the OpenAI speech endpoint. Music models are served here too.
POSThttps://api.oxyy.ai/v1/audio/speechreturns audio bytes
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | A text-to-speech or music model id, e.g. gemini-3.1-flash-tts-preview. |
| input | string | Required | The text to speak. Up to 4,096 characters. |
| voice | string | Required | The voice to speak in. OpenAI voice names (alloy, echo, fable, onyx, nova, shimmer) work on OpenAI and Gemini models. For other models, use one of the voices listed for that model — options.voices in GET /v1/models — or, on ElevenLabs models, a 20-character ElevenLabs voice id. Music and sound models take no voice. |
| response_format | string | Optional | The audio container. One of:mp3opusaacflacwavpcm Default mp3 |
| speed | number | Optional | Playback rate, 0.25–4. Default 1 |
| instructions | string | Optional | Delivery direction for the models that accept it (e.g. gpt-4o-mini-tts). |
| stream_format | string | Optional | One of:sseaudio |
| async | boolean | Optional | Return a job instead of waiting — useful for long input. Poll GET /v1/audio/speech/{id}. Default false |
| user | string | Optional | A stable id for your own end user. |
Code examples
import os from pathlib import Path from openai import OpenAI client = OpenAI( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai/v1" ) with client.audio.speech.with_streaming_response.create( model="gemini-3.1-flash-tts-preview", voice="alloy", input="Hello! This is a text-to-speech test.", response_format="mp3", speed=1.0, ) as response: response.stream_to_file(Path("output.mp3"))
# The body is raw audio bytes — write it straight to a file. curl https://api.oxyy.ai/v1/audio/speech \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{ "model": "gemini-3.1-flash-tts-preview", "voice": "alloy", "input": "Hello from Oxyy.", "response_format": "mp3" }' \ --output speech.mp3
Response
audio/mpeg
// The response body is raw audio bytes, not JSON. // // HTTP/1.1 200 OK // Content-Type: audio/mpeg // Content-Length: 45321 // X-Request-Id: 8f2c... // // Write it straight to a file or pipe it to a player. // response_format picks the container: mp3, opus, aac, flac, wav, pcm.
Available models
27 models
Model
Model ID
