File inputs & uploads
Image editing, image-to-video and transcription all take a file. Every one of them accepts the same three methods, so you can pick whichever fits your architecture.
| Method | Content-Type | Field | Best for |
|---|---|---|---|
| Upload | multipart/form-data | image / file | A file you already have on disk. The official SDKs do this for you. |
| URL | application/json | image / file_url | A file you already host, or one uploaded to the asset store. |
| Base64 | application/json | image_base64 / file_base64 | Bytes you hold in memory, with no hosting step. |
images_base64 for base64, or images for data URLs, and a multipart request may carry up to ten files. The first is the primary reference; the rest are additional context for models that blend them.The asset store
Upload once, reference many times. Useful when the same image feeds several requests, or when you would rather not resend megabytes of base64 on every call.
# 1. Upload once — the URL is reusable for 12 hours. curl https://api.oxyy.ai/v1/assets/upload \ -H "Authorization: Bearer $OXYY_API_KEY" \ -F "file=@source.png" # => {"asset_id":"...","url":"https://.../source.png","expires_at":...} # 2. Reference that URL from any media request. curl https://api.oxyy.ai/v1/images/edits \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{ "model": "gpt-image-2", "image": "https://api.oxyy.ai/storage/images/source.png", "prompt": "Change the background to a beach sunset" }'
import os, requests headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" } with open("source.png", "rb") as f: asset = requests.post( "https://api.oxyy.ai/v1/assets/upload", headers=headers, files={"file": ("source.png", f, "image/png")}, ).json() print(asset["asset_id"], asset["url"], asset["expires_at"])
{
"asset_id": "a1b2c3d4e5f6",
"url": "https://api.oxyy.ai/storage/images/a1b2c3d4e5f6.png",
"expires_at": 1700043200,
"size": 184320,
"mime_type": "image/png"
}Limits & accepted types
| Category | Maximum per file | Accepted types |
|---|---|---|
| Image | 20 MB | png, jpeg, gif, webp |
| Audio | 50 MB | mp3, wav, ogg, flac, m4a, mp4 |
| Video | 100 MB | mp4, webm, mov, avi |
| Document | 100 MB | PDF and other documents, as a chat content part |
A multipart request may carry up to 10 files and 100 MB per file overall; the per-category caps above are what each endpoint enforces, and an operator can set a lower cap on an individual model. A file_url on transcription is fetched with a 25 MB ceiling. Uploads are stored under a type derived from the content type, never from the filename you send — so scriptable formats such as SVG and HTML are refused outright.
Image editing
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Required | An image model that takes an image as input, e.g. gpt-image-2. |
| prompt | string | Required | The edit to make. |
| image | file|string | Required* | The source image — an upload, a URL or a data URL. |
| image_base64 | string|array | Required* | Base64 source image(s). |
| images_base64 | array | Optional | Several base64 references, in order. |
| mask | file | Optional | A mask marking the region to change, on models that take one. |
| n | integer | Optional | Default 1 |
| size | string | Optional | As on image generation. |
| response_format | string | Optional | One of:urlb64_json |
image, image_base64 or a multipart upload is required.# Image edit — multipart upload straight from disk. import os, requests headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" } response = requests.post( "https://api.oxyy.ai/v1/images/edits", headers=headers, files={"image": open("source.png", "rb")}, data={ "model": "gpt-image-2", "prompt": "Change the background to a beach sunset", "n": "1", "size": "1024x1024", }, ) print(response.json()["data"][0]["url"])
import fs from 'fs'; import OpenAI from 'openai'; const client = new OpenAI({ apiKey: process.env.OXYY_API_KEY, baseURL: 'https://api.oxyy.ai/v1' }); const response = await client.images.edit({ model: 'gpt-image-2', image: fs.createReadStream('source.png'), prompt: 'Change the background to a beach sunset', n: 1, size: '1024x1024', }); console.log(response.data[0].url);
curl https://api.oxyy.ai/v1/images/edits \ -H "Authorization: Bearer $OXYY_API_KEY" \ -F "model=gpt-image-2" \ -F "image=@source.png" \ -F "prompt=Change the background to a beach sunset" \ -F "n=1" \ -F "size=1024x1024"
# Base64 — no upload step, no hosting. Note the two spellings: # `image_base64` for one reference, `images_base64` for several. curl https://api.oxyy.ai/v1/images/edits \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OXYY_API_KEY" \ -d '{ "model": "gpt-image-2", "prompt": "Blend these two scenes", "images_base64": ["iVBORw0KGgo...", "data:image/png;base64,iVBORw0..."] }'
{
"created": 1700000000,
"data": [
{ "url": "https://api.oxyy.ai/storage/images/2026/01/edited_abc123.png" }
],
"usage": { "input_tokens": 9, "output_tokens": 1290, "cost": 0.04 }
}Image to video
Same parameters as video generation, with the image required rather than optional. It is a job like every other video request — poll it the same way. /v1/videos/generations also accepts an image and behaves identically.
# Image-to-video — the starting frame plus a motion prompt. import os, requests headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" } response = requests.post( "https://api.oxyy.ai/v1/videos/image-to-video", headers=headers, files={"image": open("photo.jpg", "rb")}, data={ "model": "grok-imagine-video", "prompt": "Slowly pan across the scene", "duration": "8", "resolution": "1080p", }, ) job = response.json() print(job["id"], job["status"])
curl https://api.oxyy.ai/v1/videos/image-to-video \ -H "Authorization: Bearer $OXYY_API_KEY" \ -F "model=grok-imagine-video" \ -F "image=@photo.jpg" \ -F "prompt=Slowly pan across the scene" \ -F "duration=8" \ -F "resolution=1080p"
Audio transcription
See Speech to text for the full parameter list. All three input methods apply.
