The smallest model in DeepSeek's new architecture family, with native multimodal visual understanding. Designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.
AI models
Compare pricingFast, high-quality everyday image generation.
OpenAI's most capable model for image generation and editing.
Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Meta's multimodal reasoning model for long-horizon agentic and coding workflows, with a 1M-token context window, reliable tool calling, and native understanding of video, images and documents.
High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps and custom vocabulary biasing.
Low-latency bidirectional streaming speech-to-text over WebSockets using the Live API, with interim and finalized transcription events, Smart transcription mode and multiple voice-activity-detection strategies.
Z.ai's flagship model for complex software engineering and long-horizon agent tasks. Uses the same base model as GLM-5.2, with all gains driven by post-training.
DeepSeek's flagship model, with a 1M-token context window and both thinking and non-thinking modes. Text only.
xAI's flagship model for code and everything else: agentic tool calling, minimal hallucinations, configurable reasoning.
Open-weight multimodal model distilled from Muse Spark, built to run on your own hardware. 30 billion parameters, designed for local agentic and coding workloads on a single 24 GB consumer GPU.
Coding-optimized Muse Spark release, purpose-built for agentic workflows with improvements to code generation, debugging and codebase understanding.
The Muse Spark release that opened the Meta Model API to developers. Multimodal reasoning model built for agentic tasks.
Meta's agentic image generation and editing model. Reasons through prompts, can run built-in web and image search during generation, and refines images across turns.
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image gene…
Long-context reasoning and coding model. The GLM release that expanded the context window from 200K to 1M tokens.
Previous-generation Fable-tier model for demanding reasoning and long-horizon agentic work. Superseded by Claude Fable 5.1.
Flagship model for the most complex professional work, coding and reasoning.
OpenAI's image generation and editing model with flexible, near-arbitrary output resolutions up to 4K and automatic high-fidelity handling of reference images.
Powerful, low-latency speech generation with natural outputs, steerable prompts and expressive inline audio tags for precise narration control across 70+ languages.
Encoder-free multimodal Gemma 4 model. Instead of separate vision and audio encoders, it projects raw image patches and audio waveforms straight into the LLM's embedding space through lightweight linear layers, so every…
Gemma 4's Mixture-of-Experts model. 25.2B total parameters but only 3.8B active per token, so it runs almost as fast as a 4B model while scoring close to the dense 31B.
The largest Gemma 4 model: a 30.7B-parameter dense multimodal model for reasoning, agentic workflows, coding and multimodal understanding, deployable on consumer GPUs and workstations.
The smallest Gemma 4 model, built for efficient on-device execution on phones and laptops, with native audio input.
On-device Gemma 4 model for laptops and mobile devices, with native audio input. 'E' stands for effective parameters.
High-efficiency production-scale image generation and editing, balancing speed with 4K generation, world knowledge and reliable text rendering. The generalist workhorse of the Nano Banana family.
Legacy Sonnet model. The first Sonnet with the full 1M token context window at standard pricing.
Z.ai's new-generation flagship foundation model for agentic engineering, targeting complex system engineering and long-range agent tasks.
Legacy Opus model. The first Opus with the full 1M token context window at standard pricing.
ElevenLabs' most emotionally rich, expressive speech synthesis model. Natural, life-like speech with high emotional range and contextual understanding, and support for natural multi-speaker dialogue.
Third-generation Pro model built for multimodal understanding, agentic capability, and vibe-coding; improved thinking, token efficiency, and factual grounding over Gemini 3 Pro.
Fastest, cheapest FLUX.2 model for real-time and high-volume text-to-image and multi-reference editing.
Quality-focused Klein model pairing a 9B flow model with an 8B Qwen3 text embedder, step-distilled to four steps.
Legacy first-generation Gemini 3 Flash model providing baseline speed and intelligence with frontier-class multimodal understanding.
32B-parameter open-weight FLUX.2 model for text-to-image generation and multi-reference editing.
Legacy Opus model with extended thinking and a 200K context window.
The fastest Claude model with near-frontier intelligence.
Legacy Sonnet model with a 200K token context window and extended thinking.
Smallest and most cost-effective multimodal model in the 2.5 family, built for at-scale usage.
Google's first hybrid reasoning model with configurable thinking budgets; best price-performance for low-latency, high-volume tasks that require reasoning.
Most advanced model of the 2.5 family, with deep reasoning and coding capability for complex tasks.
Fast and controllable text-to-speech for low-latency, cost-efficient applications and real-time assistants, with fine control over style and pacing.
High-fidelity speech synthesis optimized for quality in structured workflows such as podcasts and audiobooks, with more natural outputs and easier-to-steer prompts.
FP8-quantized checkpoint of Qwen3-30B-A3B, a Mixture-of-Experts model with hybrid thinking and non-thinking modes.
Text-to-speech model optimized for realtime, low-latency use.
Text-to-speech model optimized for audio quality. Double the price of tts-1.
The model currently powering ChatGPT, exposed through the API. A rolling alias whose underlying snapshot changes over time.
Anthropic's most capable model, for demanding reasoning and long-horizon agentic work.
Legacy Opus model with adaptive thinking and the xhigh effort level for long-running agentic and coding tasks.
Previous-generation Opus model with adaptive thinking. Anthropic recommends migrating to Claude Opus 5.
For complex agentic coding and enterprise work. A step-change over Claude Opus 4.8 in deep reasoning, agentic and long-horizon tasks, and test-time compute scaling.
The best combination of speed and intelligence in the Claude lineup.
Z.ai's video generation model.
Z.ai's CogView image generation model.
Dubs video with automatic speaker detection across 29 languages.
End-to-end dubbing model that preserves voice and emotion across 92 languages.
English-only voice changer model.
English-only ultra-fast speech synthesis model.
ElevenLabs' fastest speech synthesis model, built for real-time applications and the Agents Platform. Balances speed and quality at half the price per character.
Lifelike, consistent-quality speech synthesis. The most stable model for long-form generation, with consistent voice quality and accent across language switches.
State-of-the-art multilingual voice changer. Converts an existing recording into another voice while preserving delivery.
State-of-the-art multilingual voice designer model.
Studio-grade music generation from text prompts, composition plans and previously generated songs.
ElevenLabs' most advanced music model. Studio-grade generation from text prompts, composition plans and previously generated songs, with richer melodies, deeper arrangements and more layered instruments than music_v2.
Human-like and expressive voice DESIGN model. Generates a new voice from a text description rather than speech from text.
ElevenLabs' most expressive realtime speech synthesis model, tuned for low-latency conversation while keeping v3's emotional range.
ElevenLabs' speech models combined in one low-latency agent pipeline, for adding voice to a chat agent.
Cost-efficient multimodal model for high-volume agentic tasks, translation, and simple data extraction where budget and latency are the primary constraints.
Legacy Flash model providing sustained frontier-level intelligence for real-world tasks; effective for sub-agent deployment, multi-step workflows, and long-horizon tasks at scale.
Fastest, most cost-effective model in the 3.5 family, optimized for high-throughput agentic tasks, translation, and simple data processing.
Previous-generation Flash model balancing speed and multimodal capability across general agentic and everyday tasks; strong at code generation, agentic execution, and spatial reasoning.
High-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution.
Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into a unified embedding space for semantic search and RAG.
Fast conversational video generation and editing with native audio, keyframe interpolation and clip extension. Turn text and images into video and refine results through natural language.
Z.ai's translation agent.
Z.ai agent that generates slides and posters.
32-billion-parameter GLM-4 model with a 128K context window, flat-priced on input and output.
GLM-4.5 base text model. The release that introduced interleaved reasoning to the GLM line.
Lightweight, low-cost variant of GLM-4.5.
High-throughput variant of GLM-4.5-Air.
Free tier text model from the GLM-4.5 generation.
Highest-performance variant of GLM-4.5, and the most expensive model in Z.ai's published catalog.
Previous-generation GLM vision-language model.
Prior-generation GLM-4 text model.
Multimodal model for high-fidelity visual understanding and long-context reasoning across images, documents and mixed media. Handles complex page layouts and charts as visual input.
Free tier vision-language model.
High-throughput, low-cost vision variant of GLM-4.6V.
Previous flagship focused on task completion rather than single-point code generation, with interleaved, retained and round-level reasoning.
Free tier text model from the GLM-4.7 generation.
High-throughput, very low cost variant of GLM-4.7.
Coding-specialized variant of GLM-5.
Speed-optimized variant of GLM-5.
Leading open-source coding model with significant gains on long-horizon tasks. Predecessor to GLM-5.2.
The first native multimodal model in the GLM-5 series, delivering stronger intelligence than GLM-5.2 at very low cost.
Higher-throughput FlashX variant of the GLM-5.3 generation.
Z.ai's automatic speech recognition model.
Text-to-image generation model, described by Z.ai as achieving open-source state of the art in complex scenarios.
Z.ai's OCR model for extracting text from images and documents.
Versatile, high-intelligence GPT model. Accepts text and image input and produces text, including Structured Outputs.
Fast, affordable small model for focused tasks. Text and image input, text output.
Text-to-speech model powered by GPT-4o Mini, with an instructions parameter for tone and style steering.
GPT-5.1 reasoning model with a no-reasoning default for fast responses.
Previous flagship model for complex professional work.
OpenAI's Codex model, optimized for agentic coding workflows.
OpenAI's strongest mini model for coding, computer use and subagents.
OpenAI's most advanced cybersecurity model, for authorized vulnerability research and security testing.
GPT-5.6 model optimized for cost-sensitive, high-volume workloads.
Flagship model for complex professional work.
GPT-5.6 model that balances intelligence and cost.
OpenAI's most capable model, built for the hardest end-to-end work.
OpenAI's premier model for natural, expressive voice conversations with smooth interruption handling.
Low-latency speech-to-text model for realtime transcription. The recommended replacement for whisper-1 in live transcription.
Voice model for audio in, audio out.
Realtime reasoning model with tool use. Previous generation to GPT-Realtime-2.1.
Realtime reasoning model with tool use for speech-to-speech agents.
Cost-efficient realtime reasoning model with tool use.
Streaming speech-to-speech translation model.
Streaming speech-to-text model for realtime transcription.
High-accuracy speech-to-text model for file and Realtime input transcription. The recommended replacement for whisper-1 file transcription.
Pinned 0309 release of Grok 4.20 configured for multi-agent orchestration.
Pinned 0309 release of Grok 4.20 in non-reasoning mode, for low-latency responses.
Pinned 0309 release of Grok 4.20 in reasoning mode, with a 1M token context window.
Cost-efficient Grok model with a 1M token context window and a 20% Batch API discount.
Previous-generation Grok flagship. Same standard input and output rates as Grok 4.6 but a cheaper cached-input rate.
xAI's coding-specialized model, the cheapest per token in the Grok lineup.
xAI's low-cost image generation and editing model, flat-priced across resolutions.
xAI's recommended image generation and editing model. Text and image in, image out.
Previous-generation, lower-cost video generation and editing model. Text, image and video in, video out.
xAI's recommended video generation model. Text, image and audio in, video out, with native audio.
xAI's speech-to-text transcription, available as batch REST or streaming.
xAI's text-to-speech mode, with custom voice support.
xAI's realtime speech-to-speech voice agent model.
Muse Spark 1.2 at discounted token pricing, in exchange for permission for Meta to train future models on your data.
Muse Spark 1.3 at heavily discounted token pricing, in exchange for permission for Meta to train future models on your prompts and completions.
Meta's speech-to-text model for streaming and file transcription, with speaker attribution and turn detection built into the model.
Z.ai agent that applies popular special-effects templates to video.
Open-weight segmentation model served on Meta Model API. Name an object in a short text prompt and get a box and pixel-accurate mask for every match; in video it follows each object across frames.
State-of-the-art batch speech recognition across 90+ languages, with precise word-level timestamps, speaker diarization and dynamic audio tagging.
A fine-tune of Scribe v2 for medical and clinical audio. Improves recognition of drug names, anatomy, pathology and clinical dictation while matching Scribe v2 on everyday speech.
ElevenLabs' fastest and most accurate live speech recognition model, delivering partial transcriptions in about 150ms across 90+ languages.
Royalty-free sound effects generation from text prompts.
Removes background noise, ambient sounds, reverb and interference to leave clean dialogue.
142 of 142 models












