Responses API
Responses API
The OpenAI Responses API, for clients built on client.responses.create such as the OpenAI Agents SDK. It runs on the same models, routing and billing as chat completions, and every chat model is available through it. Use the OpenAI SDK with base_url set to https://api.oxyy.ai/v1.
POSThttps://api.oxyy.ai/v1/responses
Stateless. Responses are not stored. Send the full conversation in
input on every turn. previous_response_id is not supported and is refused with a 400 that says what to send instead. The hosted tools that need a service this gateway does not run (file_search, code_interpreter,computer_use, mcp) are refused the same way; web_search andimage_generation are served.Code examples
# pip install openai import os from openai import OpenAI client = OpenAI( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai/v1" ) response = client.responses.create( model="gpt-5.5", instructions="Answer in one sentence.", input="What is the capital of France?" ) print(response.output_text)
// npm install openai import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.OXYY_API_KEY, baseURL: 'https://api.oxyy.ai/v1' }); const response = await client.responses.create({ model: 'gpt-5.5', instructions: 'Answer in one sentence.', input: 'What is the capital of France?' }); console.log(response.output_text);
curl https://api.oxyy.ai/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{"model": "gpt-5.5", "instructions": "Answer in one sentence.", "input": "What is the capital of France?"}'
Example response
Response
{
"id": "resp_abc123def456",
"object": "response",
"created_at": 1700000000,
"status": "completed",
"model": "gpt-5.5",
"output": [
{
"id": "msg_abc123def456",
"type": "message",
"status": "completed",
"role": "assistant",
"content": [
{ "type": "output_text", "text": "The capital of France is Paris.", "annotations": [] }
]
}
],
"output_text": "The capital of France is Paris.",
"previous_response_id": null,
"store": false,
"usage": {
"input_tokens": 21,
"input_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
"output_tokens": 8,
"output_tokens_details": { "reasoning_tokens": 0 },
"total_tokens": 29,
"cost": 0.000106
}
}Multi-turn conversations
Keep the conversation in your application: append the previous output items and the next user message to input, then send all of it again. Reasoning and function-call items from output can be sent back as they are.
history = [{"role": "user", "content": "Name a prime number."}]
response = client.responses.create(model="gpt-5.5", input=history)
print(response.output_text)
# Nothing is stored between calls: append the reply and the next
# question, then send the whole conversation again.
history += [item.model_dump(exclude_none=True) for item in response.output]
history.append({"role": "user", "content": "Now double it."})
response = client.responses.create(model="gpt-5.5", input=history)
print(response.output_text)const input = [{ role: 'user', content: 'Name a prime number.' }]; let response = await client.responses.create({ model: 'gpt-5.5', input }); console.log(response.output_text); // Nothing is stored between calls: append the reply and the next // question, then send the whole conversation again. input.push(...response.output, { role: 'user', content: 'Now double it.' }); response = await client.responses.create({ model: 'gpt-5.5', input }); console.log(response.output_text);
Supported features
| Feature | Status | Notes |
|---|---|---|
input as a string or a list of items | Supported | message, function_call, function_call_output and reasoning items |
instructions, max_output_tokens, temperature, top_p | Supported | |
stream | Supported | Server-sent events such as response.output_text.delta and response.completed |
Function tools, tool_choice, parallel_tool_calls | Supported | tool_choice accepts none, auto, required, a function or allowed_tools |
text.format | Supported | text, json_object and json_schema |
reasoning.effort | Supported | |
input_image and input_file parts | Supported | Images as a URL or a data URL; files as file_data (a data URL). file_id and file_url are refused. |
Hosted tool: image_generation | Supported | Its size, quality, background and output_format drive an image model; other fields on the tool are reported as dropped. |
Hosted tool: web_search / web_search_preview | Supported | Becomes web_search_options on the chat request, so the model answers from the live web and cites its sources as url_citation annotations. search_context_size and user_location are honoured; other fields on the tool are reported as dropped. The provider bills the search on top of the tokens. |
store, include, truncation, metadata | Supported | Accepted, but nothing is stored: store is always returned as false |
previous_response_id, conversation | Not supported | Refused with a 400. Send the full conversation in input instead. |
background, stored prompt templates, item_reference items | Not supported | Refused with a 400 |
Other hosted tools: file_search, code_interpreter, computer_use, mcp, shell | Not supported | Refused with a 400. Declare a function tool and run it in your application. |
| Retrieving, deleting or cancelling a response by ID | Not supported | Only POST /v1/responses is available |
