OpenAI: gpt-oss-20b

openai/gpt-oss-20b

OpenAI gpt-oss-20b API

Oxyy has not published a price for gpt-oss-20b yet. It accepts text and returns text, with a 131K-token context window. For every OpenAI model and rate, see OpenAI API pricing.

Smaller open-weight model for lower latency, local and specialized use. Delivers results similar to OpenAI o3-mini on common benchmarks.
ModalitiesTextText
PriceNot yet publishedAll OpenAI API pricing
Context131K
ReleasedAug 5, 2025

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderPrice
OpenAINot yet published

Capabilities (9)

StreamingTool callingStructured outputReasoningFull Chain Of ThoughtFine TunableLocal InferenceSystem promptHarmony Format Required

Supported parameters (3)

reasoning_efforttoolsresponse_format

Any other parameter you send is ignored rather than rejected.

Pricing

Models on Oxyy are billed pay as you go, with no subscription. This one has no published price yet.

Price
Not yet published
Oxyy has not published a price for this model yet.
RateStandard priceChargedUnit
Price not yet published

Specifications

Model ID
gpt-oss-20b
Developer
OpenAI
Type
Text
Input
Text
Output
Text
Context window
131,072 tokens
Released
Aug 5, 2025
Tokenizer
GPT
Series
gpt-oss
Endpoint
/v1/chat/completions

Best for

Open weightsReasoningAgenticLocal inferenceApache 2.0

Quickstart

Call gpt-oss-20b with a POST to /v1/chat/completions on https://api.oxyy.ai, using your Oxyy API key from the OXYY_API_KEY environment variable. The same key works for every model in the catalog.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.oxyy.ai/v1",
    api_key=os.environ["OXYY_API_KEY"],
)

response = client.chat.completions.create(
    model="gpt-oss-20b",
    messages=[
        {"role": "user", "content": "Explain unified AI APIs in one sentence."},
    ],
)
print(response.choices[0].message.content)

Availability

The share of requests to this model that completed successfully over the last 30 days. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Requests (30 days)
3
A success rate is shown from 100 completed requests
Days with traffic
2
of the last 30 days

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt78
Reasoning12
Completion22

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length.

Frequently asked questions

What is gpt-oss-20b?
Smaller open-weight model for lower latency, local and specialized use. Delivers results similar to OpenAI o3-mini on common benchmarks.
What is the context length of gpt-oss-20b?
gpt-oss-20b has a context window of 131,072 tokens.
Does gpt-oss-20b support tool calling and structured outputs?
Yes — gpt-oss-20b supports both tool calling and structured outputs.
What inputs and outputs does gpt-oss-20b support?
gpt-oss-20b accepts text as input and returns text.
Which API endpoint does gpt-oss-20b use?
Send a POST request to https://api.oxyy.ai/v1/chat/completions with "model": "gpt-oss-20b", authenticated with your Oxyy API key.
When was gpt-oss-20b released?
gpt-oss-20b was released on Aug 5, 2025.
What other models does OpenAI have?
OpenAI also offers chat-latest, GPT-4o, GPT-4o mini and GPT-5.1 through Oxyy.

Alternatives to gpt-oss-20b

Models of the same kind from other vendors, closest in price first. Prices are pay-as-you-go on Oxyy, through the same API key.

Explore more models

More models from OpenAI