Streaming
Streaming
Set stream: true on chat completions, responses or either vendor SDK and the reply arrives as server-sent events. Every official SDK handles the framing for you; the shape below is what you get over plain HTTP.
import os from openai import OpenAI client = OpenAI( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai/v1" ) stream = client.chat.completions.create( model="gpt-5.5", messages=[{"role": "user", "content": "Write a haiku about caching."}], stream=True, stream_options={"include_usage": True}, # final chunk carries usage + cost ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) elif chunk.usage: print(f"\n\ncost: {chunk.usage.cost}")
import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.OXYY_API_KEY, baseURL: 'https://api.oxyy.ai/v1' }); const stream = await client.chat.completions.create({ model: 'gpt-5.5', messages: [{ role: 'user', content: 'Write a haiku about caching.' }], stream: true, stream_options: { include_usage: true }, }); for await (const chunk of stream) { const delta = chunk.choices?.[0]?.delta?.content; if (delta) process.stdout.write(delta); if (chunk.usage) console.log('\ncost:', chunk.usage.cost); }
curl -N https://api.oxyy.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{ "model": "gpt-5.5", "messages": [{"role": "user", "content": "Hello!"}], "stream": true, "stream_options": {"include_usage": true} }'
On the wire
text/event-stream
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{"content":"Cold"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
// Only when stream_options.include_usage is true — choices is empty here.
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":17,"total_tokens":29,"cost":0.000106}}
data: [DONE]| Parameter | Type | Required | Description |
|---|---|---|---|
| stream | boolean | Optional | Stream the reply as server-sent events. Default false |
| stream_options.include_usage | boolean | Optional | Emit one final chunk carrying usage — token counts and cost — with an empty choices array, just before [DONE]. Default false |
Things worth knowing. The stream always ends with
data: [DONE]. Closing the connection cancels the upstream generation, and you are billed only for what was produced. A failure that happens after the first byte arrives as an error frame inside the stream rather than an HTTP status, because the status was already sent.Idempotency-Key is ignored on streaming requests — a stream cannot be replayed.