Audio API

Speech to text

Transcribe audio with timestamps, language detection and speaker diarization, in several output formats.

POSThttps://api.oxyy.ai/v1/audio/transcriptionstranscribe in the source language
POSThttps://api.oxyy.ai/v1/audio/translationstranscribe into English

Parameters

ParameterTypeRequiredDescription
modelstringRequiredA speech-to-text model id, e.g. scribe_v2, whisper-large-v3-turbo.
filefileRequired*The audio, as multipart/form-data. Containers: mp3, mp4, mpeg, mpga, m4a, wav, webm, ogg, flac.
file_urlstringRequired*A public URL to fetch instead of uploading. Up to 25 MB; private and loopback addresses are refused.
file_base64stringRequired*Base64 audio, with or without a data: prefix.
response_formatstringOptionalOne of:jsontextverbose_jsonsrtvttdiarized_json Default json
languagestringOptionalISO-639-1 code (en, ja, bn). Omit to auto-detect.
promptstringOptionalVocabulary or style hint — names, jargon, expected spelling.
temperaturenumberOptionalSampling temperature, 0–1.
timestamp_granularitiesarrayOptionalRequires verbose_json. One of:wordsegment
diarizebooleanOptionalLabel speakers, on the models that support it.
num_speakersintegerOptionalA hint for diarization.
*Exactly one of file, file_url or file_base64 is required. /v1/audio/translations takes the same parameters and returns English text. Not every transcription model can translate; one that cannot answers with a 400 that says so.

Code examples

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OXYY_API_KEY"],
    base_url="https://api.oxyy.ai/v1"
)

# Supported containers: mp3, mp4, mpeg, mpga, m4a, wav, webm, ogg, flac
with open("interview.mp3", "rb") as audio_file:
    response = client.audio.transcriptions.create(
        model="scribe_v2",
        file=audio_file,
        response_format="verbose_json",
        timestamp_granularities=["segment"],
        language="en",
    )

print(response.text)
for seg in response.segments:
    print(f"{seg.start:.1f}s-{seg.end:.1f}s: {seg.text}")
import fs from 'fs';
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.OXYY_API_KEY,
  baseURL: 'https://api.oxyy.ai/v1'
});

const response = await client.audio.transcriptions.create({
  model: 'scribe_v2',
  file: fs.createReadStream('interview.mp3'),
  response_format: 'verbose_json',
  language: 'en',
});

console.log(response.text);
# multipart upload
curl https://api.oxyy.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -F "model=scribe_v2" \
  -F "file=@interview.mp3" \
  -F "response_format=verbose_json" \
  -F "language=en"

# …or JSON with a hosted URL instead of an upload
curl https://api.oxyy.ai/v1/audio/transcriptions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -d '{"model": "scribe_v2", "file_url": "https://example.com/interview.mp3"}'
# Same inputs as /transcriptions; the transcript comes back in English.
curl https://api.oxyy.ai/v1/audio/translations \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -F "model=whisper-large-v3-turbo" \
  -F "file=@bengali-interview.mp3" \
  -F "response_format=json"

Example response (verbose_json)

Response
{
  "task": "transcribe",
  "language": "en",
  "duration": 5.42,
  "text": "Hello, this is a transcription test.",
  "segments": [
    { "id": 0, "start": 0.0, "end": 2.1, "text": "Hello, this is" },
    { "id": 1, "start": 2.1, "end": 5.42, "text": "a transcription test." }
  ]
}

verbose_json is requested from the provider as sent. The one exception is the text-shaped formats: text, srt and vtt are rendered here from a timed result, so they work on every provider.

Available models