oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Models
Filters
4K64K1M
FREE$0.5$10+
FREE$2$30+
New12+ mo
0%100%

AI models

Compare pricing
DeepSeek-V4.1-FlashTextImage→Text-60%14d ago
reasoningcodinglong-contextvisionlow-cost+1 more

The smallest model in DeepSeek's new architecture family, with native multimodal visual understanding. Designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.

by DeepSeekSep 10, 20261M context$0.150/M input$0.600/M output
OpenAI: GPT-Image-2.5 FlareTextImage→Image-60%16d ago
image-generationimage-editinglow-latency

Fast, high-quality everyday image generation.

by OpenAISep 8, 2026— context$5.00/M input—/M output
OpenAI: GPT-Image-2.5 SunburstTextImage→Image-60%16d ago
image-generationimage-editing

OpenAI's most capable model for image generation and editing.

by OpenAISep 8, 2026— context$5.00/M input—/M output
Google: Gemini 3.8 FlashTextImageVideoAudioFile→Text-60%320.9M tokens

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.

by GoogleSep 2, 20261.05M context$0.750/M input$3.75/M output7019ms latency734 t/s
Meta: Muse Spark 1.3TextImageVideoFile→Text-60%22d ago
reasoningcodingagenticlong-contextmultimodal

Meta's multimodal reasoning model for long-horizon agentic and coding workflows, with a 1M-token context window, reliable tool calling, and native understanding of video, images and documents.

by MetaSep 2, 20261.05M context$1.25/M input$4.25/M output
Google: Gemini 3.5 TranscribeAudioText→Text-60%29d ago

High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps and custom vocabulary biasing.

by GoogleAug 26, 202696K context$2.00/M input$12.00/M output
Google: Gemini 3.5 Transcribe LiveAudio→Text-60%29d ago

Low-latency bidirectional streaming speech-to-text over WebSockets using the Live API, with interim and finalized transcription events, Smart transcription mode and multiple voice-activity-detection strategies.

by GoogleAug 26, 202696K context$3.50/M input$21.00/M output
Z.ai: GLM-5.3Text→Text-60%1.1M tokens
reasoningcodingagenticlong-contextopen-weights

Z.ai's flagship model for complex software engineering and long-horizon agent tasks. Uses the same base model as GLM-5.2, with all gains driven by post-training.

by Z.aiAug 14, 20261.05M context$1.40/M input$4.40/M output11.9s latency63 t/s
DeepSeek-V4-Pro-0813Text→Text-60%107 tokens
reasoningcodinglong-contextflagship

DeepSeek's flagship model, with a 1M-token context window and both thinking and non-thinking modes. Text only.

by DeepSeekAug 13, 20261M context$0.660/M input$1.98/M output3120ms latency16 t/s
xAI: Grok 4.6TextImage→Text-60%311.4K tokens
reasoningcodingagenticlong-context

xAI's flagship model for code and everything else: agentic tool calling, minimal hallucinations, configurable reasoning.

by xAIAug 12, 2026500K context$2.00/M input$6.00/M output5263ms latency48 t/s
Meta: Muse GlimmerTextImage→Text-60%1mo ago
open-weightslocal-inferencecodingagenticmultimodal

Open-weight multimodal model distilled from Muse Spark, built to run on your own hardware. 30 billion parameters, designed for local agentic and coding workloads on a single 24 GB consumer GPU.

by MetaAug 10, 2026— context—/M input—/M output
Meta: Muse Spark 1.2TextImageVideoFile→Text-60%1mo ago
codingagenticlong-contextmultimodal

Coding-optimized Muse Spark release, purpose-built for agentic workflows with improvements to code generation, debugging and codebase understanding.

by MetaAug 5, 20261.05M context$1.25/M input$4.25/M output
Meta: Muse Spark 1.1TextImageVideoFile→Text-60%2mo ago
reasoningagenticlong-contextmultimodal

The Muse Spark release that opened the Meta Model API to developers. Multimodal reasoning model built for agentic tasks.

by MetaJul 9, 20261.05M context$1.25/M input$4.25/M output
Meta: Muse Image 1.0TextImage→Image-60%2mo ago
image-generationimage-editingagentic

Meta's agentic image generation and editing model. Reasons through prompts, can run built-in web and image search during generation, and refines images across turns.

by MetaJul 7, 2026— context—/M input—/M output
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)ImageText→ImageText-60%127.0M tokens

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image gene…

by GoogleJun 30, 202666K context$0.250/M input$1.50/M output4557ms latency
Z.ai: GLM-5.2Text→Text-60%1.8B tokens
reasoningcodingagenticlong-contextopen-weights

Long-context reasoning and coding model. The GLM release that expanded the context window from 200K to 1M tokens.

by Z.aiJun 16, 20261.05M context$1.40/M input$4.40/M output19.3s latency70 t/s
Anthropic: Claude Fable 5TextImage→Text-60%140.6K tokens
reasoningcodingagenticlong-context

Previous-generation Fable-tier model for demanding reasoning and long-horizon agentic work. Superseded by Claude Fable 5.1.

by AnthropicJun 9, 20261M context$10.00/M input$50.00/M output2938ms latency52 t/s
OpenAI: GPT-5.5TextImage→Text-60%69.0K tokens

Flagship model for the most complex professional work, coding and reasoning.

by OpenAIApr 24, 20261.05M context$5.00/M input$30.00/M output6932ms latency6 t/s
OpenAI: GPT Image 2TextImage→Image-60%294.7K tokens
image-generationimage-editinginpainting4ktext-rendering

OpenAI's image generation and editing model with flexible, near-arbitrary output resolutions up to 4K and automatic high-fidelity handling of reference images.

by OpenAIApr 21, 2026— context$2.50/M input—/M output21.4s latency
Google: Gemini 3.1 Flash TTSText→Audio-60%16.5K tokens

Powerful, low-latency speech generation with natural outputs, steerable prompts and expressive inline audio tags for precise narration control across 70+ languages.

by GoogleApr 15, 202633K context$1.00/M input—/M output5434ms latency
Google: Gemma 4 12B UnifiedTextImageAudioVideo→Text-60%5mo ago
open-weightsmultimodalaudioencoder-freelocal-inference

Encoder-free multimodal Gemma 4 model. Instead of separate vision and audio encoders, it projects raw image patches and audio waveforms straight into the LLM's embedding space through lightweight linear layers, so every…

by GoogleApr 2, 2026262K context$0.100/M input$0.300/M output
Google: Gemma 4 26B A4B Instruct (MoE)TextImageVideo→Text-60%32.8M tokens
open-weightsmoereasoningcodingfast+1 more

Gemma 4's Mixture-of-Experts model. 25.2B total parameters but only 3.8B active per token, so it runs almost as fast as a 4B model while scoring close to the dense 31B.

by GoogleApr 2, 2026262K context$0.130/M input$0.400/M output8572ms latency60 t/s
Google: Gemma 4 31B InstructTextImageVideo→Text-60%2.2M tokens
open-weightsreasoningcodingagenticmultimodal+1 more

The largest Gemma 4 model: a 30.7B-parameter dense multimodal model for reasoning, agentic workflows, coding and multimodal understanding, deployable on consumer GPUs and workstations.

by GoogleApr 2, 2026262K context$0.270/M input$0.760/M output48.5s latency36 t/s
Google: Gemma 4 E2BTextImageAudioVideo→Text-60%5mo ago
open-weightson-devicemobileaudioedge+1 more

The smallest Gemma 4 model, built for efficient on-device execution on phones and laptops, with native audio input.

by GoogleApr 2, 2026131K context$0.040/M input$0.040/M output
Google: Gemma 4 E4BTextImageAudioVideo→Text-60%5mo ago
open-weightson-devicemobileaudioedge

On-device Gemma 4 model for laptops and mobile devices, with native audio input. 'E' stands for effective parameters.

by GoogleApr 2, 2026131K context$0.200/M input$0.200/M output
Google: Nano Banana 2TextImageVideo→TextImage-60%176.7M tokens

High-efficiency production-scale image generation and editing, balancing speed with 4K generation, world knowledge and reliable text rendering. The generalist workhorse of the Nano Banana family.

by GoogleFeb 26, 2026— context$0.500/M input$3.00/M output9949ms latency
Anthropic: Claude Sonnet 4.6TextImage→Text-60%33.1M tokens
reasoningcodingagenticlong-context

Legacy Sonnet model. The first Sonnet with the full 1M token context window at standard pricing.

by AnthropicFeb 17, 20261M context$3.00/M input$15.00/M output14.0s latency83 t/s
Z.ai: GLM-5Text→Text-60%51.3K tokens
reasoningcodingagenticopen-weights

Z.ai's new-generation flagship foundation model for agentic engineering, targeting complex system engineering and long-range agent tasks.

by Z.aiFeb 11, 2026200K context$1.00/M input$3.20/M output2558ms latency443 t/s
Anthropic: Claude Opus 4.6TextImage→Text-60%6.3M tokens
reasoningcodingagenticlong-context

Legacy Opus model. The first Opus with the full 1M token context window at standard pricing.

by AnthropicFeb 5, 20261M context$5.00/M input$25.00/M output2785ms latency94 t/s
EElevenLabs: Eleven v3Text→Audio-60%7mo ago
text-to-speechexpressiveaudiobookmulti-speakeraudio-tags

ElevenLabs' most emotionally rich, expressive speech synthesis model. Natural, life-like speech with high emotional range and contextual understanding, and support for natural multi-speaker dialogue.

by ElevenLabsFeb 1, 2026— context—/M input—/M output5085ms latency
Google: Gemini 3.1 ProTextImageVideoAudioFile→Text-60%124.1M tokens

Third-generation Pro model built for multimodal understanding, agentic capability, and vibe-coding; improved thinking, token efficiency, and factual grounding over Gemini 3 Pro.

by GoogleFeb 1, 20261.05M context$2.00/M input$12.00/M output13.8s latency257 t/s
BBlack Forest Labs: FLUX.2 [klein] 4BTextImage→Image-60%8mo ago

Fastest, cheapest FLUX.2 model for real-time and high-volume text-to-image and multi-reference editing.

by Black Forest LabsJan 15, 2026— context—/M input—/M output0ms latency
BBlack Forest Labs: FLUX.2 [klein] 9BTextImage→Image-60%8mo ago

Quality-focused Klein model pairing a 9B flow model with an 8B Qwen3 text embedder, step-distilled to four steps.

by Black Forest LabsJan 15, 2026— context—/M input—/M output0ms latency
Google: Gemini 3 FlashTextImageVideoAudioFile→Text-60%3.9B tokens

Legacy first-generation Gemini 3 Flash model providing baseline speed and intelligence with frontier-class multimodal understanding.

by GoogleJan 1, 20261.05M context$0.500/M input$3.00/M output12.5s latency885 t/s
BBlack Forest Labs: FLUX.2 [dev]TextImage→Image-60%10mo ago

32B-parameter open-weight FLUX.2 model for text-to-image generation and multi-reference editing.

by Black Forest LabsNov 25, 2025— context—/M input—/M output
Anthropic: Claude Opus 4.5TextImage→Text-60%61.6K tokens
reasoningcodingagentic

Legacy Opus model with extended thinking and a 200K context window.

by AnthropicNov 24, 2025200K context$5.00/M input$25.00/M output1330ms latency95 t/s
Anthropic: Claude Haiku 4.5TextImage→Text-60%8.2M tokens
fastlow-costclassificationhigh-volume

The fastest Claude model with near-frontier intelligence.

by AnthropicOct 1, 2025200K context$1.00/M input$5.00/M output1021ms latency102 t/s
Anthropic: Claude Sonnet 4.5TextImageFile→Text-60%684.4K tokens
reasoningcodingagentic

Legacy Sonnet model with a 200K token context window and extended thinking.

by AnthropicSep 29, 2025200K context$3.00/M input$15.00/M output1496ms latency94 t/s
Google: Gemini 2.5 Flash-LiteTextImageVideoAudioFile→Text-60%177.6M tokens

Smallest and most cost-effective multimodal model in the 2.5 family, built for at-scale usage.

by GoogleJul 22, 20251.05M context$0.100/M input$0.400/M output2068ms latency218 t/s
Google: Gemini 2.5 FlashTextImageVideoAudioFile→Text-60%121.3M tokens

Google's first hybrid reasoning model with configurable thinking budgets; best price-performance for low-latency, high-volume tasks that require reasoning.

by GoogleJun 17, 20251.05M context$0.300/M input$2.50/M output2097ms latency230 t/s
Google: Gemini 2.5 ProTextImageVideoAudioFile→Text-60%1.5M tokens

Most advanced model of the 2.5 family, with deep reasoning and coding capability for complex tasks.

by GoogleJun 17, 20251.05M context$1.25/M input$10.00/M output30.2s latency140 t/s
Google: Gemini 2.5 Flash TTSText→Audio-60%45 tokens

Fast and controllable text-to-speech for low-latency, cost-efficient applications and real-time assistants, with fine control over style and pacing.

by GoogleMay 20, 20258K context$0.500/M input—/M output4241ms latency
Google: Gemini 2.5 Pro TTSText→Audio-60%1y ago

High-fidelity speech synthesis optimized for quality in structured workflows such as podcasts and audiobooks, with more natural outputs and easier-to-steer prompts.

by GoogleMay 20, 20258K context$1.00/M input—/M output
Qwen (Alibaba): Qwen3 30B A3B (FP8)Text→Text-60%4.3K tokens

FP8-quantized checkpoint of Qwen3-30B-A3B, a Mixture-of-Experts model with hybrid thinking and non-thinking modes.

by Qwen (Alibaba)Apr 28, 202533K context—/M input—/M output1394ms latency165 t/s
OpenAI: TTS-1Text→Audio-60%2y ago

Text-to-speech model optimized for realtime, low-latency use.

by OpenAINov 6, 2023— context—/M input—/M output2126ms latency
OpenAI: TTS-1 HDText→Audio-60%2y ago

Text-to-speech model optimized for audio quality. Double the price of tts-1.

by OpenAINov 6, 2023— context—/M input—/M output2750ms latency
OpenAI: chat-latestTextImage→Text-60%—
chatrolling-alias

The model currently powering ChatGPT, exposed through the API. A rolling alias whose underlying snapshot changes over time.

by OpenAI— context$5.00/M input$30.00/M output
Anthropic: Claude Fable 5.1TextImage→Text-60%892.8K tokens
reasoningcodingagenticlong-context

Anthropic's most capable model, for demanding reasoning and long-horizon agentic work.

by Anthropic1M context$10.00/M input$50.00/M output4015ms latency45 t/s
Anthropic: Claude Opus 4.7TextImage→Text-60%1.8M tokens
reasoningcodingagenticlong-context

Legacy Opus model with adaptive thinking and the xhigh effort level for long-running agentic and coding tasks.

by Anthropic1M context$5.00/M input$25.00/M output1415ms latency99 t/s
Anthropic: Claude Opus 4.8TextImage→Text-60%21.9M tokens
reasoningcodingagenticlong-context

Previous-generation Opus model with adaptive thinking. Anthropic recommends migrating to Claude Opus 5.

by Anthropic1M context$5.00/M input$25.00/M output6576ms latency70 t/s
Anthropic: Claude Opus 5TextImage→Text-60%283.9M tokens
reasoningcodingagenticlong-context

For complex agentic coding and enterprise work. A step-change over Claude Opus 4.8 in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.

by Anthropic1M context$5.00/M input$25.00/M output51.6s latency74 t/s
Anthropic: Claude Sonnet 5TextImage→Text-60%151.1M tokens
reasoningcodingagenticlong-context

The best combination of speed and intelligence in the Claude lineup.

by Anthropic1M context$2.00/M input$10.00/M output29.4s latency85 t/s
Z.ai: CogVideoX-3TextImage→Video-60%—
video-generationopen-weights

Z.ai's video generation model.

by Z.ai— context—/M input—/M output
Z.ai: CogView-4Text→Image-60%—
image-generationtext-to-imageopen-weightslow-cost

Z.ai's CogView image generation model.

by Z.ai— context—/M input—/M output
EElevenLabs: Dubbing v1AudioVideo→AudioVideo-60%—
dubbingvideospeaker-detection

Dubs video with automatic speaker detection across 29 languages.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Dubbing v2AudioVideo→AudioVideo-60%—
dubbingvideovoice-preservation

End-to-end dubbing model that preserves voice and emotion across 92 languages.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven English v2 (Voice Changer)Audio→Audio-60%—
voice-changerspeech-to-speechenglish-only

English-only voice changer model.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Flash v2Text→Audio-60%—
text-to-speechlow-latencyenglish-only

English-only ultra-fast speech synthesis model.

by ElevenLabs— context—/M input—/M output513ms latency
EElevenLabs: Eleven Flash v2.5Text→Audio-60%—
text-to-speechlow-latencyrealtimebulkvoice-agents

ElevenLabs' fastest speech synthesis model, built for real-time applications and the Agents Platform. Balances speed and quality at half the price per character.

by ElevenLabs— context—/M input—/M output430ms latency
EElevenLabs: Eleven Multilingual v2Text→Audio-60%—
text-to-speechlong-formstablenarration

Lifelike, consistent-quality speech synthesis. The most stable model for long-form generation, with consistent voice quality and accent across language switches.

by ElevenLabs— context—/M input—/M output1961ms latency
EElevenLabs: Eleven Multilingual v2 (Voice Changer)Audio→Audio-60%—
voice-changerspeech-to-speechmultilingual

State-of-the-art multilingual voice changer. Converts an existing recording into another voice while preserving delivery.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Multilingual v2 (Voice Design)Text→Audio-60%—
voice-designtext-to-voicemultilingual

State-of-the-art multilingual voice designer model.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Music v2TextAudio→Audio-60%—
music-generationtext-to-music

Studio-grade music generation from text prompts, composition plans and previously generated songs.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven Music v2.5TextAudio→Audio-60%—
music-generationtext-to-musicstudio-grade

ElevenLabs' most advanced music model. Studio-grade generation from text prompts, composition plans and previously generated songs, with richer melodies, deeper arrangements and more layered instruments than music_v2.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven v3 (Voice Design / Text to Voice)Text→Audio-60%—
voice-designtext-to-voice

Human-like and expressive voice DESIGN model. Generates a new voice from a text description rather than speech from text.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Eleven v3 ConversationalText→Audio-60%—
text-to-speechrealtimeexpressivevoice-agentsaudio-tags

ElevenLabs' most expressive realtime speech synthesis model, tuned for low-latency conversation while keeping v3's emotional range.

by ElevenLabs— context—/M input—/M output
EElevenLabs: ElevenAgents Speech EngineTextAudio→Audio-60%—
voice-agentsrealtimespeech-to-speechpipeline

ElevenLabs' speech models combined in one low-latency agent pipeline, for adding voice to a chat agent.

by ElevenLabs— context—/M input—/M output
Google: Gemini 3.1 Flash-LiteTextImageVideoAudioFile→Text-60%226.4M tokens

Cost-efficient multimodal model for high-volume agentic tasks, translation, and simple data extraction where budget and latency are the primary constraints.

by Google1.05M context$0.250/M input$1.50/M output1772ms latency264 t/s
Google: Gemini 3.5 FlashTextImageVideoAudioFile→Text-60%14.1M tokens

Legacy Flash model providing sustained frontier-level intelligence for real-world tasks; effective for sub-agent deployment, multi-step workflows, and long-horizon tasks at scale.

by Google1.05M context$1.50/M input$9.00/M output12.3s latency317 t/s
Google: Gemini 3.5 Flash-LiteTextImageVideoAudioFile→Text-60%59.1M tokens

Fastest, most cost-effective model in the 3.5 family, optimized for high-throughput agentic tasks, translation, and simple data processing.

by Google1.05M context$0.300/M input$2.50/M output2355ms latency268 t/s
Google: Gemini 3.6 FlashTextImageVideoAudioFile→Text-60%18.0M tokens

Previous-generation Flash model balancing speed and multimodal capability across general agentic and everyday tasks; strong at code generation, agentic execution, and spatial reasoning.

by Google1.05M context$0.750/M input$3.75/M output3269ms latency742 t/s
Google: Gemini 3.7 FlashTextImageVideoAudioFile→Text-60%151.2M tokens

High-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution.

by Google1.05M context$0.750/M input$3.75/M output11.6s latency515 t/s
Google: Gemini Embedding 2TextImageAudioVideoFile→Embeddings-60%—
embeddingmultimodalragsemantic-search

Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into a unified embedding space for semantic search and RAG.

by Google— context$0.200/M input—/M output
Google: Gemini Omni FlashTextImageVideoAudio→TextVideoAudio-60%—

Fast conversational video generation and editing with native audio, keyframe interpolation and clip extension. Turn text and images into video and refine results through natural language.

by Google— context$1.50/M input$9.00/M output
Z.ai: General-Purpose TranslationText→Text-60%—
agenttranslation

Z.ai's translation agent.

by Z.ai— context$3.00/M input—/M output
Z.ai: GLM Slide/Poster Agent (beta)Text→TextImage-60%—
agentslidesposterbeta

Z.ai agent that generates slides and posters.

by Z.ai— context$0.700/M input—/M output
Z.ai: GLM-4-32B-0414-128KText→Text-60%—
low-costopen-weightslong-context

32-billion-parameter GLM-4 model with a 128K context window, flat-priced on input and output.

by Z.ai128K context$0.100/M input$0.100/M output
Z.ai: GLM-4.5Text→Text-60%—
codingagenticopen-weights

GLM-4.5 base text model. The release that introduced interleaved reasoning to the GLM line.

by Z.ai— context$0.600/M input$2.20/M output
Z.ai: GLM-4.5-AirText→Text-60%—
low-costopen-weights

Lightweight, low-cost variant of GLM-4.5.

by Z.ai— context$0.200/M input$1.10/M output
Z.ai: GLM-4.5-AirXText→Text-60%—
high-throughput

High-throughput variant of GLM-4.5-Air.

by Z.ai— context$1.10/M input$4.50/M output
Z.ai: GLM-4.5-FlashText→TextFree—
freeopen-weights

Free tier text model from the GLM-4.5 generation.

by Z.ai— context$0/M input$0/M output
Z.ai: GLM-4.5-XText→Text-60%—
reasoninghigh-performance

Highest-performance variant of GLM-4.5, and the most expensive model in Z.ai's published catalog.

by Z.ai— context$2.20/M input$8.90/M output
Z.ai: GLM-4.5VTextImage→Text-60%—
visionmultimodalopen-weights

Previous-generation GLM vision-language model.

by Z.ai— context$0.600/M input$1.80/M output
Z.ai: GLM-4.6Text→Text-60%—
codingagenticopen-weights

Prior-generation GLM-4 text model.

by Z.ai— context$0.600/M input$2.20/M output
Z.ai: GLM-4.6VTextImage→Text-60%—
visionmultimodaldocument-understandingui-reconstruction

Multimodal model for high-fidelity visual understanding and long-context reasoning across images, documents and mixed media. Handles complex page layouts and charts as visual input.

by Z.ai128K context$0.300/M input$0.900/M output
Z.ai: GLM-4.6V-FlashTextImage→TextFree—
visionfree

Free tier vision-language model.

by Z.ai— context$0/M input$0/M output
Z.ai: GLM-4.6V-FlashXTextImage→Text-60%—
visionlow-costhigh-throughput

High-throughput, low-cost vision variant of GLM-4.6V.

by Z.ai— context$0.040/M input$0.400/M output
Z.ai: GLM-4.7Text→Text-60%—
reasoningcodingagenticopen-weights

Previous flagship focused on task completion rather than single-point code generation, with interleaved, retained and round-level reasoning.

by Z.ai200K context$0.600/M input$2.20/M output
Z.ai: GLM-4.7-FlashText→TextFree83.4K tokens
freeopen-weights

Free tier text model from the GLM-4.7 generation.

by Z.ai— context$0/M input$0/M output6507ms latency65 t/s
Z.ai: GLM-4.7-FlashXText→Text-60%—
low-costhigh-throughput

High-throughput, very low cost variant of GLM-4.7.

by Z.ai— context$0.070/M input$0.400/M output
Z.ai: GLM-5-CodeText→Text-60%—
codingagentic

Coding-specialized variant of GLM-5.

by Z.ai— context$1.20/M input$5.00/M output
Z.ai: GLM-5-TurboText→Text-60%—
reasoningcodingfast

Speed-optimized variant of GLM-5.

by Z.ai— context$1.20/M input$4.00/M output
Z.ai: GLM-5.1Text→Text-60%51.2K tokens
reasoningcodingagenticopen-weights

Leading open-source coding model with significant gains on long-horizon tasks. Predecessor to GLM-5.2.

by Z.ai200K context$1.40/M input$4.40/M output4566ms latency53 t/s
Z.ai: GLM-5.3-FlashTextImage→Text-60%281.7K tokens
codingmultimodallow-costopen-weightsvisual-coding

The first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 at very low cost.

by Z.ai— context$0.150/M input$0.500/M output14.8s latency63 t/s
Z.ai: GLM-5.3-FlashXText→Text-60%—
codinglow-costhigh-throughput

Higher-throughput FlashX variant of the GLM-5.3 generation.

by Z.ai— context$0.370/M input$1.25/M output
Z.ai: GLM-ASR-2512Audio→Text-60%—
speech-to-texttranscriptionlow-cost

Z.ai's automatic speech recognition model.

by Z.ai— context$0.030/M input—/M output
Z.ai: GLM-ImageText→Image-60%—
image-generationtext-to-imageopen-weights

Text-to-image generation model, described by Z.ai as achieving open-source state of the art in complex scenarios.

by Z.ai— context—/M input—/M output
Z.ai: GLM-OCRImageFile→Text-60%—
ocrdocument-understandinglow-cost

Z.ai's OCR model for extracting text from images and documents.

by Z.ai— context$0.030/M input$0.030/M output
OpenAI: GPT-4oTextImage→Text-60%244.1K tokens

Versatile, high-intelligence GPT model. Accepts text and image input and produces text, including Structured Outputs.

by OpenAI128K context$2.50/M input$10.00/M output5144ms latency59 t/s
OpenAI: GPT-4o miniTextImage→Text-60%1.3M tokens

Fast, affordable small model for focused tasks. Text and image input, text output.

by OpenAI128K context$0.150/M input$0.600/M output1432ms latency32 t/s
OpenAI: GPT-4o Mini TTSText→Audio-60%—
text-to-speechsteerable

Text-to-speech model powered by GPT-4o Mini, with an instructions parameter for tone and style steering.

by OpenAI— context—/M input—/M output1659ms latency
OpenAI: GPT-5.1TextImage→Text-60%675.1K tokens

GPT-5.1 reasoning model with a no-reasoning default for fast responses.

by OpenAI400K context$1.25/M input$10.00/M output5479ms latency89 t/s
OpenAI: GPT-5.2TextImage→Text-60%159.0K tokens

Previous flagship model for complex professional work.

by OpenAI400K context$1.75/M input$14.00/M output3089ms latency47 t/s
OpenAI: GPT-5.3 CodexTextImage→Text-60%—
codingagentic

OpenAI's Codex model, optimized for agentic coding workflows.

by OpenAI— context$1.75/M input$14.00/M output
OpenAI: GPT-5.4 miniTextImage→Text-60%1.1K tokens

OpenAI's strongest mini model for coding, computer use and subagents.

by OpenAI400K context$0.750/M input$4.50/M output1147ms latency43 t/s
OpenAI: GPT-5.6 Cyber (Daybreak Red)TextImage→Text-60%—
cybersecurityvulnerability-researchrestricted-access

OpenAI's most advanced cybersecurity model, for authorized vulnerability research and security testing.

by OpenAI— context$12.50/M input$75.00/M output
OpenAI: GPT-5.6 LunaTextImage→Text-60%863 tokens
low-costhigh-volumelong-context

GPT-5.6 model optimized for cost-sensitive, high-volume workloads.

by OpenAI1.05M context$0.200/M input$1.20/M output2741ms latency55 t/s
OpenAI: GPT-5.6 SolTextImage→Text-60%2.3M tokens
reasoningcodingagenticlong-context

Flagship model for complex professional work.

by OpenAI1.05M context$4.00/M input$20.00/M output8371ms latency32 t/s
OpenAI: GPT-5.6 TerraTextImage→Text-60%2.7K tokens
reasoningcodingbalancedlong-context

GPT-5.6 model that balances intelligence and cost.

by OpenAI1.05M context$2.00/M input$12.00/M output3456ms latency48 t/s
OpenAI: GPT-6 AstraTextImage→Text-60%217 tokens
reasoningcodingagenticlong-context

OpenAI's most capable model, built for the hardest end-to-end work.

by OpenAI1.05M context$10.00/M input$50.00/M output2606ms latency33 t/s
OpenAI: GPT-Live 1TextAudio→TextAudio-60%—
voicerealtimespeech-to-speech

OpenAI's premier model for natural, expressive voice conversations with smooth interruption handling.

by OpenAI— context—/M input—/M output
OpenAI: GPT-Live-TranscribeAudio→Text-60%—
speech-to-textrealtimelow-latency

Low-latency speech-to-text model for realtime transcription. The recommended replacement for whisper-1 in live transcription.

by OpenAI— context—/M input—/M output
OpenAI: GPT-Realtime-1.5TextAudio→TextAudio-60%—
realtimevoice

Voice model for audio in, audio out.

by OpenAI— context—/M input—/M output
OpenAI: GPT-Realtime-2TextAudio→TextAudio-60%—
realtimevoice

Realtime reasoning model with tool use. Previous generation to GPT-Realtime-2.1.

by OpenAI— context—/M input—/M output
OpenAI: GPT-Realtime-2.1TextAudioImage→TextAudio-60%—
realtimevoicespeech-to-speech

Realtime reasoning model with tool use for speech-to-speech agents.

by OpenAI— context$4.00/M input$24.00/M output
OpenAI: GPT-Realtime-2.1 MiniTextAudioImage→TextAudio-60%—
realtimevoicelow-cost

Cost-efficient realtime reasoning model with tool use.

by OpenAI— context$0.600/M input$2.40/M output
OpenAI: GPT-Realtime-TranslateAudio→AudioText-60%—
translationrealtimespeech-to-speech

Streaming speech-to-speech translation model.

by OpenAI— context—/M input—/M output
OpenAI: GPT-Realtime-WhisperAudio→Text-60%—
speech-to-textrealtimestreaming

Streaming speech-to-text model for realtime transcription.

by OpenAI— context—/M input—/M output
OpenAI: GPT-TranscribeAudio→Text-60%—
speech-to-texttranscription

High-accuracy speech-to-text model for file and Realtime input transcription. The recommended replacement for whisper-1 file transcription.

by OpenAI— context—/M input—/M output
xAI: Grok 4.20 (0309) Multi-AgentTextImage→Text-60%1.2M tokens
multi-agentreasoninglong-contextpinned-snapshot

Pinned 0309 release of Grok 4.20 configured for multi-agent orchestration.

by xAI1M context$1.25/M input$2.50/M output5820ms latency8737 t/s
xAI: Grok 4.20 (0309) Non-ReasoningTextImage→Text-60%45.4K tokens
long-contextlow-latencypinned-snapshot

Pinned 0309 release of Grok 4.20 in non-reasoning mode, for low-latency responses.

by xAI1M context$1.25/M input$2.50/M output671ms latency121 t/s
xAI: Grok 4.20 (0309) ReasoningTextImage→Text-60%111.3K tokens
reasoninglong-contextpinned-snapshot

Pinned 0309 release of Grok 4.20 in reasoning mode, with a 1M token context window.

by xAI1M context$1.25/M input$2.50/M output3420ms latency342 t/s
xAI: Grok 4.3TextImage→Text-60%94.5K tokens
reasoninglong-contextcost-efficient

Cost-efficient Grok model with a 1M token context window and a 20% Batch API discount.

by xAI1M context$1.25/M input$2.50/M output2444ms latency206 t/s
xAI: Grok 4.5TextImage→Text-60%14.0M tokens
reasoningcodingagenticlong-context

Previous-generation Grok flagship. Same standard input and output rates as Grok 4.6 but a cheaper cached-input rate.

by xAI500K context$2.00/M input$6.00/M output19.9s latency50 t/s
xAI: Grok Build 0.1TextImage→Text-60%109.1K tokens
codingagenticlow-cost

xAI's coding-specialized model, the cheapest per token in the Grok lineup.

by xAI256K context$1.00/M input$2.00/M output3874ms latency79 t/s
xAI: Grok Imagine ImageTextImage→Image-60%—
image-generationimage-editinglow-cost

xAI's low-cost image generation and editing model, flat-priced across resolutions.

by xAI— context—/M input—/M output
xAI: Grok Imagine Image 2.0TextImage→Image-60%—
image-generationimage-editing

xAI's recommended image generation and editing model. Text and image in, image out.

by xAI— context—/M input—/M output
xAI: Grok Imagine VideoTextImageVideo→Video-60%—
video-generationvideo-editinglow-cost

Previous-generation, lower-cost video generation and editing model. Text, image and video in, video out.

by xAI— context—/M input—/M output
xAI: Grok Imagine Video 1.5TextImageAudio→VideoAudio-60%—
video-generationimage-to-videoaudio

xAI's recommended video generation model. Text, image and audio in, video out, with native audio.

by xAI— context—/M input—/M output
xAI: Grok Voice API — Speech to TextAudio→Text-60%—
speech-to-texttranscriptionstreaming

xAI's speech-to-text transcription, available as batch REST or streaming.

by xAI— context—/M input—/M output
xAI: Grok Voice API — Text to SpeechText→Audio-60%—
text-to-speechcustom-voices

xAI's text-to-speech mode, with custom voice support.

by xAI— context—/M input—/M output
xAI: Grok Voice Think Fast 2.0TextAudio→TextAudio-60%—
voicerealtimespeech-to-speech

xAI's realtime speech-to-speech voice agent model.

by xAI— context—/M input—/M output
Meta: Muse Spark 1.2 (Contributor tier)TextImageVideoFile→Text-60%—
codingagenticlow-costtrains-on-your-data

Muse Spark 1.2 at discounted token pricing, in exchange for permission for Meta to train future models on your data.

by Meta1.05M context$0.100/M input$0.200/M output
Meta: Muse Spark 1.3 (Contributor tier)TextImageVideoFile→Text-60%—
reasoningcodingagenticlow-costtrains-on-your-data

Muse Spark 1.3 at heavily discounted token pricing, in exchange for permission for Meta to train future models on your prompts and completions.

by Meta1.05M context$0.100/M input$0.200/M output
Meta: Muse Voice Transcribe 1.0Audio→Text-60%—
speech-to-texttranscriptiondiarizationrealtime

Meta's speech-to-text model for streaming and file transcription, with speaker attribution and turn detection built into the model.

by Meta— context—/M input—/M output
Z.ai: Popular Special Effects Video TemplatesTextImageVideo→Video-60%—
agentvideoeffects

Z.ai agent that applies popular special-effects templates to video.

by Z.ai— context—/M input—/M output
Meta: SAM 3.1TextImageVideo→Embeddings-60%—
segmentationvisionvideoopen-weights

Open-weight segmentation model served on Meta Model API. Name an object in a short text prompt and get a box and pixel-accurate mask for every match; in video it follows each object across frames.

by Meta— context—/M input—/M output
EElevenLabs: Scribe v2Audio→Text-60%—
speech-to-textdiarizationtimestampsentity-detection

State-of-the-art batch speech recognition across 90+ languages, with precise word-level timestamps, speaker diarization and dynamic audio tagging.

by ElevenLabs— context—/M input—/M output1055ms latency
EElevenLabs: Scribe v2 MedicalAudio→Text-60%—
speech-to-textmedicalclinicalhipaadiarization

A fine-tune of Scribe v2 for medical and clinical audio. Improves recognition of drug names, anatomy, pathology and clinical dictation while matching Scribe v2 on everyday speech.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Scribe v2 RealtimeAudio→Text-60%—
speech-to-textrealtimestreaminglow-latencyvoice-agents

ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Sound Effects v2Text→Audio-60%—
sound-effectsaudio-generationroyalty-free

Royalty-free sound effects generation from text prompts.

by ElevenLabs— context—/M input—/M output
EElevenLabs: Voice IsolatorAudio→Audio-60%—
audio-processingnoise-removal

Removes background noise, ambient sounds, reverb and interference to leave clean dialogue.

by ElevenLabs— context—/M input—/M output

142 of 142 models

oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout