oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. xAI/
  4. Grok 4.20 (0309) Non-Reasoning

xAI: Grok 4.20 (0309) Non-Reasoning

xai/grok-4.20-0309-non-reasoning
Compare Playground Get API key
Pinned 0309 release of Grok 4.20 in non-reasoning mode, for low-latency responses.
ModalitiesTextImageText
In / out price$0.500 / $1.00 per 1M
Context1M
Released—
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
xAI$0.500$1.00$0.2000.79s121 tps

Capabilities (11)

StreamingTool callingStructured outputSystem promptVisionPrompt cachingWeb searchX SearchCode executionPriority ProcessingBatch

Supported parameters (2)

toolsresponse_format

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
$0.500
/M tokens · listed $1.25
Effective output price
$1.00
/M tokens · listed $2.50
RateListedChargedUnit
Input$1.25$0.500per 1M tokens
Output$2.50$1.00per 1M tokens
Cache read$0.200$0.200per 1M tokens

Long-context pricing. Once the prompt is larger than the threshold, the whole request is billed at these rates per 1M tokens; rates not listed keep the price above.From 200,000 tokens in the prompt: input $2.50 / output $5.00 / cache read $0.400.

UTC

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
121tok/s
P50, streaming requests
Latency
0.79s
P50, end to end
UTC

Throughput

P992184 tok/s
P951560 tok/s
P50101 tok/s

Latency

P506990ms
P757061ms
P957598ms
P998641ms

Time to first token

P50550ms
P95874ms

Cache hit rate

51.74%
Share of input tokens served from the prompt cache over the window.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
91.50%
over 118 requests
Days with traffic
2
of the last 30

Days are UTC days. Days with no requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

  1. 1.Oxyy Tester44.2K tokens

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt43.3K
Completion2.1K

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is Grok 4.20 (0309) Non-Reasoning?
Pinned 0309 release of Grok 4.20 in non-reasoning mode, for low-latency responses.
How much does Grok 4.20 (0309) Non-Reasoning cost?
Grok 4.20 (0309) Non-Reasoning costs $1.25 per million input tokens and $2.50 per million output tokens. Cached prompt tokens are billed at $0.200 per million. You pay per request, with no subscription.
What is the context length of Grok 4.20 (0309) Non-Reasoning?
Grok 4.20 (0309) Non-Reasoning has a context window of 1,000,000 tokens.
Does Grok 4.20 (0309) Non-Reasoning support tool calling and structured outputs?
Yes — Grok 4.20 (0309) Non-Reasoning supports both tool calling and structured outputs.
What inputs and outputs does Grok 4.20 (0309) Non-Reasoning support?
Grok 4.20 (0309) Non-Reasoning accepts text, image as input and returns text.
What other models does xAI have?
xAI also offers Grok 4.20 (0309) Reasoning, Grok 4.20 (0309) Multi-Agent, Grok 4.3, Grok 4.5 through Oxyy.
When was Grok 4.20 (0309) Non-Reasoning released?
No release date is published for Grok 4.20 (0309) Non-Reasoning.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from xAI

Grok 4.20 (0309) ReasoningPinned 0309 release of Grok 4.20 in reasoning mode, with a 1M token context window.text · 1M context · $1.25 / $2.50Grok 4.20 (0309) Multi-AgentPinned 0309 release of Grok 4.20 configured for multi-agent orchestration.text · 1M context · $1.25 / $2.50Grok 4.3Cost-efficient Grok model with a 1M token context window and a 20% Batch API discount.text · 1M context · $1.25 / $2.50Grok 4.5Previous-generation Grok flagship. Same standard input and output rates as Grok 4.6 but a cheaper cached-input rate.text · 500K context · $2.00 / $6.00Grok 4.6xAI's flagship model for code and everything else: agentic tool calling, minimal hallucinations, configurable reasoning.text · 500K context · $2.00 / $6.00Grok Build 0.1xAI's coding-specialized model, the cheapest per token in the Grok lineup.text · 256K context · $1.00 / $2.00
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout