oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. xAI/
  4. Grok Voice API — Speech to Text

xAI: Grok Voice API — Speech to Text

xai/grok-stt
Compare Playground Get API key
xAI's speech-to-text transcription, available as batch REST or streaming.
ModalitiesAudioText
In / out price— / — per 1M
Context—
Released—
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
xAI—————

Capabilities (1)

Streaming

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
—
/M tokens
Effective output price
—
/M tokens
RateListedChargedUnit
UTC
Price history starts building from the first daily snapshot — there is not enough of it yet to draw a trend.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
—tok/s
P50, streaming requests
Latency
—s
P50, end to end
UTC
No requests to this model in the selected window, so there is nothing to measure yet.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
—
over 0 requests
Days with traffic
0
of the last 30
This model has not served any requests in the last 30 days, so there is no availability to report.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

No apps have identified themselves for this model yet. Clients can opt in by sending an X-Title header with their request.

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

No traffic in this window
Prompt0
Completion0

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is Grok Voice API — Speech to Text?
xAI's speech-to-text transcription, available as batch REST or streaming.
How much does Grok Voice API — Speech to Text cost?
Pricing for Grok Voice API — Speech to Text is listed on this page under Pricing. You pay per request, with no subscription.
What is the context length of Grok Voice API — Speech to Text?
The context window for Grok Voice API — Speech to Text is not published.
Does Grok Voice API — Speech to Text support tool calling and structured outputs?
Neither tool calling nor structured outputs are listed among this model's capabilities. Any unsupported parameter you send is ignored rather than rejected.
What inputs and outputs does Grok Voice API — Speech to Text support?
Grok Voice API — Speech to Text accepts audio as input and returns text.
What other models does xAI have?
xAI also offers Grok 4.20 (0309) Non-Reasoning, Grok 4.20 (0309) Reasoning, Grok 4.20 (0309) Multi-Agent, Grok 4.3 through Oxyy.
When was Grok Voice API — Speech to Text released?
No release date is published for Grok Voice API — Speech to Text.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from xAI

Grok 4.20 (0309) Non-ReasoningPinned 0309 release of Grok 4.20 in non-reasoning mode, for low-latency responses.text · 1M context · $1.25 / $2.50Grok 4.20 (0309) ReasoningPinned 0309 release of Grok 4.20 in reasoning mode, with a 1M token context window.text · 1M context · $1.25 / $2.50Grok 4.20 (0309) Multi-AgentPinned 0309 release of Grok 4.20 configured for multi-agent orchestration.text · 1M context · $1.25 / $2.50Grok 4.3Cost-efficient Grok model with a 1M token context window and a 20% Batch API discount.text · 1M context · $1.25 / $2.50Grok 4.5Previous-generation Grok flagship. Same standard input and output rates as Grok 4.6 but a cheaper cached-input rate.text · 500K context · $2.00 / $6.00Grok 4.6xAI's flagship model for code and everything else: agentic tool calling, minimal hallucinations, configurable reasoning.text · 500K context · $2.00 / $6.00
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout