Streaming

Streaming

Set stream: true on chat completions, responses or either vendor SDK and the reply arrives as server-sent events. Every official SDK handles the framing for you; the shape below is what you get over plain HTTP.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OXYY_API_KEY"],
    base_url="https://api.oxyy.ai/v1"
)

stream = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Write a haiku about caching."}],
    stream=True,
    stream_options={"include_usage": True},  # final chunk carries usage + cost
)

for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    elif chunk.usage:
        print(f"\n\ncost: {chunk.usage.cost}")
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.OXYY_API_KEY,
  baseURL: 'https://api.oxyy.ai/v1'
});

const stream = await client.chat.completions.create({
  model: 'gpt-5.5',
  messages: [{ role: 'user', content: 'Write a haiku about caching.' }],
  stream: true,
  stream_options: { include_usage: true },
});

for await (const chunk of stream) {
  const delta = chunk.choices?.[0]?.delta?.content;
  if (delta) process.stdout.write(delta);
  if (chunk.usage) console.log('\ncost:', chunk.usage.cost);
}
curl -N https://api.oxyy.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -d '{
    "model": "gpt-5.5",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true,
    "stream_options": {"include_usage": true}
  }'

On the wire

text/event-stream
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{"content":"Cold"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

// Only when stream_options.include_usage is true — choices is empty here.
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1700000000,"model":"gpt-5.5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":17,"total_tokens":29,"cost":0.000106}}

data: [DONE]
ParameterTypeRequiredDescription
streambooleanOptionalStream the reply as server-sent events. Default false
stream_options.include_usagebooleanOptionalEmit one final chunk carrying usage — token counts and cost — with an empty choices array, just before [DONE]. Default false
Things worth knowing. The stream always ends with data: [DONE]. Closing the connection cancels the upstream generation, and you are billed only for what was produced. A failure that happens after the first byte arrives as an error frame inside the stream rather than an HTTP status, because the status was already sent.Idempotency-Key is ignored on streaming requests — a stream cannot be replayed.