File inputs

File inputs & uploads

Image editing, image-to-video and transcription all take a file. Every one of them accepts the same three methods, so you can pick whichever fits your architecture.

MethodContent-TypeFieldBest for
Uploadmultipart/form-dataimage / fileA file you already have on disk. The official SDKs do this for you.
URLapplication/jsonimage / file_urlA file you already host, or one uploaded to the asset store.
Base64application/jsonimage_base64 / file_base64Bytes you hold in memory, with no hosting step.
Several references at once. Image endpoints accept an array: images_base64 for base64, or images for data URLs, and a multipart request may carry up to ten files. The first is the primary reference; the rest are additional context for models that blend them.

The asset store

Upload once, reference many times. Useful when the same image feeds several requests, or when you would rather not resend megabytes of base64 on every call.

POSThttps://api.oxyy.ai/v1/assets/uploadmultipart, field: file
GEThttps://api.oxyy.ai/v1/assets/{id}metadata
GEThttps://api.oxyy.ai/v1/assetsyour assets
DELETEhttps://api.oxyy.ai/v1/assets/{id}delete
# 1. Upload once — the URL is reusable for 12 hours.
curl https://api.oxyy.ai/v1/assets/upload \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -F "file=@source.png"

# => {"asset_id":"...","url":"https://.../source.png","expires_at":...}

# 2. Reference that URL from any media request.
curl https://api.oxyy.ai/v1/images/edits \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -d '{
    "model": "gpt-image-2",
    "image": "https://api.oxyy.ai/storage/images/source.png",
    "prompt": "Change the background to a beach sunset"
  }'
import os, requests

headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" }

with open("source.png", "rb") as f:
    asset = requests.post(
        "https://api.oxyy.ai/v1/assets/upload",
        headers=headers,
        files={"file": ("source.png", f, "image/png")},
    ).json()

print(asset["asset_id"], asset["url"], asset["expires_at"])
Response
{
  "asset_id": "a1b2c3d4e5f6",
  "url": "https://api.oxyy.ai/storage/images/a1b2c3d4e5f6.png",
  "expires_at": 1700043200,
  "size": 184320,
  "mime_type": "image/png"
}
Assets expire after 12 hours. The store is a staging area for requests in flight, not file hosting. Anything you need to keep, keep yourself.

Limits & accepted types

CategoryMaximum per fileAccepted types
Image20 MBpng, jpeg, gif, webp
Audio50 MBmp3, wav, ogg, flac, m4a, mp4
Video100 MBmp4, webm, mov, avi
Document100 MBPDF and other documents, as a chat content part

A multipart request may carry up to 10 files and 100 MB per file overall; the per-category caps above are what each endpoint enforces, and an operator can set a lower cap on an individual model. A file_url on transcription is fetched with a 25 MB ceiling. Uploads are stored under a type derived from the content type, never from the filename you send — so scriptable formats such as SVG and HTML are refused outright.

Image editing

POSThttps://api.oxyy.ai/v1/images/edits
ParameterTypeRequiredDescription
modelstringRequiredAn image model that takes an image as input, e.g. gpt-image-2.
promptstringRequiredThe edit to make.
imagefile|stringRequired*The source image — an upload, a URL or a data URL.
image_base64string|arrayRequired*Base64 source image(s).
images_base64arrayOptionalSeveral base64 references, in order.
maskfileOptionalA mask marking the region to change, on models that take one.
nintegerOptionalDefault 1
sizestringOptionalAs on image generation.
response_formatstringOptionalOne of:urlb64_json
*One of image, image_base64 or a multipart upload is required.
# Image edit — multipart upload straight from disk.
import os, requests

headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" }

response = requests.post(
    "https://api.oxyy.ai/v1/images/edits",
    headers=headers,
    files={"image": open("source.png", "rb")},
    data={
        "model": "gpt-image-2",
        "prompt": "Change the background to a beach sunset",
        "n": "1",
        "size": "1024x1024",
    },
)
print(response.json()["data"][0]["url"])
import fs from 'fs';
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.OXYY_API_KEY,
  baseURL: 'https://api.oxyy.ai/v1'
});

const response = await client.images.edit({
  model: 'gpt-image-2',
  image: fs.createReadStream('source.png'),
  prompt: 'Change the background to a beach sunset',
  n: 1,
  size: '1024x1024',
});

console.log(response.data[0].url);
curl https://api.oxyy.ai/v1/images/edits \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -F "model=gpt-image-2" \
  -F "image=@source.png" \
  -F "prompt=Change the background to a beach sunset" \
  -F "n=1" \
  -F "size=1024x1024"
# Base64 — no upload step, no hosting. Note the two spellings:
# `image_base64` for one reference, `images_base64` for several.
curl https://api.oxyy.ai/v1/images/edits \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "Blend these two scenes",
    "images_base64": ["iVBORw0KGgo...", "data:image/png;base64,iVBORw0..."]
  }'
Response
{
  "created": 1700000000,
  "data": [
    { "url": "https://api.oxyy.ai/storage/images/2026/01/edited_abc123.png" }
  ],
  "usage": { "input_tokens": 9, "output_tokens": 1290, "cost": 0.04 }
}

Image to video

POSThttps://api.oxyy.ai/v1/videos/image-to-video

Same parameters as video generation, with the image required rather than optional. It is a job like every other video request — poll it the same way. /v1/videos/generations also accepts an image and behaves identically.

# Image-to-video — the starting frame plus a motion prompt.
import os, requests

headers = { "Authorization": f"Bearer {os.environ['OXYY_API_KEY']}" }

response = requests.post(
    "https://api.oxyy.ai/v1/videos/image-to-video",
    headers=headers,
    files={"image": open("photo.jpg", "rb")},
    data={
        "model": "grok-imagine-video",
        "prompt": "Slowly pan across the scene",
        "duration": "8",
        "resolution": "1080p",
    },
)
job = response.json()
print(job["id"], job["status"])
curl https://api.oxyy.ai/v1/videos/image-to-video \
  -H "Authorization: Bearer $OXYY_API_KEY" \
  -F "model=grok-imagine-video" \
  -F "image=@photo.jpg" \
  -F "prompt=Slowly pan across the scene" \
  -F "duration=8" \
  -F "resolution=1080p"

Audio transcription

See Speech to text for the full parameter list. All three input methods apply.