Multimodal input
Images, audio & documents as input
Send an array of content parts instead of a string, and a model that accepts that modality will read them. Both the OpenAI spelling and the Responses spelling are understood on /v1/chat/completions:
| Part | Carries | Source |
|---|---|---|
image_url / input_image | An image | {"url": "https://…"} or a data:image/…;base64,… URL |
input_audio | An audio clip | {"data": "<base64>", "format": "wav"} |
file / input_file | A PDF or other document | {"filename": "…", "file_data": "data:application/pdf;base64,…"} |
import base64, os from openai import OpenAI client = OpenAI( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai/v1" ) # A public https URL, or a data URL you build yourself. with open("chart.png", "rb") as f: b64 = base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="gemini-3-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What does this chart show?"}, {"type": "image_url", "image_url": { "url": f"data:image/png;base64,{b64}", "detail": "auto", }}, ], }], ) print(response.choices[0].message.content)
curl https://api.oxyy.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{ "model": "gemini-3-flash", "messages": [{ "role": "user", "content": [ {"type": "text", "text": "Describe this image."}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}} ] }] }'
# A PDF or other document rides in a `file` content part. import base64, os from openai import OpenAI client = OpenAI( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai/v1" ) with open("contract.pdf", "rb") as f: b64 = base64.b64encode(f.read()).decode() response = client.chat.completions.create( model="gemini-3-flash", messages=[{ "role": "user", "content": [ {"type": "text", "text": "Summarise the termination clause."}, {"type": "file", "file": { "filename": "contract.pdf", "file_data": f"data:application/pdf;base64,{b64}", }}, ], }], )
A model only reads what it declares. Each model lists its
inputModalities in GET /v1/models. Attaching an image to a text-only model is refused rather than quietly dropped, because a silently ignored attachment produces a confident answer about nothing. Per-file caps are 20 MB for images, 50 MB for audio and 100 MB for video and documents, and an operator can set a lower cap per model.URLs are fetched by us, not by the provider. A URL you supply is downloaded, size-checked and inlined before the provider call, so private and loopback addresses are refused. If a URL is unreachable you get a
400 that says so, not a confusing provider error.