Legacy first-generation Gemini 3 Flash model providing baseline speed and intelligence with frontier-class multimodal understanding.
Google models
Access 25 Google models through the Oxyy unified API including Gemini 3 Flash, Gemini 3.8 Flash, Gemini 3.1 Flash-Lite. Compare pricing, context windows and capabilities between Google models.
Google tokens processed on Oxyy· daily, UTC
Models 25
Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Cost-efficient multimodal model for high-volume agentic tasks, translation, and simple data extraction where budget and latency are the primary constraints.
Smallest and most cost-effective multimodal model in the 2.5 family, built for at-scale usage.
High-efficiency production-scale image generation and editing, balancing speed with 4K generation, world knowledge and reliable text rendering. The generalist workhorse of the Nano Banana family.
High-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution.
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image gene…
Third-generation Pro model built for multimodal understanding, agentic capability, and vibe-coding; improved thinking, token efficiency, and factual grounding over Gemini 3 Pro.
Google's first hybrid reasoning model with configurable thinking budgets; best price-performance for low-latency, high-volume tasks that require reasoning.
Fastest, most cost-effective model in the 3.5 family, optimized for high-throughput agentic tasks, translation, and simple data processing.
Gemma 4's Mixture-of-Experts model. 25.2B total parameters but only 3.8B active per token, so it runs almost as fast as a 4B model while scoring close to the dense 31B.
Previous-generation Flash model balancing speed and multimodal capability across general agentic and everyday tasks; strong at code generation, agentic execution, and spatial reasoning.
Legacy Flash model providing sustained frontier-level intelligence for real-world tasks; effective for sub-agent deployment, multi-step workflows, and long-horizon tasks at scale.
The largest Gemma 4 model: a 30.7B-parameter dense multimodal model for reasoning, agentic workflows, coding and multimodal understanding, deployable on consumer GPUs and workstations.
Most advanced model of the 2.5 family, with deep reasoning and coding capability for complex tasks.
Powerful, low-latency speech generation with natural outputs, steerable prompts and expressive inline audio tags for precise narration control across 70+ languages.
Fast and controllable text-to-speech for low-latency, cost-efficient applications and real-time assistants, with fine control over style and pacing.
High-accuracy, low-latency non-streaming speech-to-text with utterance-based language detection across 85+ languages, speaker diarization, word-level timestamps and custom vocabulary biasing.
Low-latency bidirectional streaming speech-to-text over WebSockets using the Live API, with interim and finalized transcription events, Smart transcription mode and multiple voice-activity-detection strategies.
Encoder-free multimodal Gemma 4 model. Instead of separate vision and audio encoders, it projects raw image patches and audio waveforms straight into the LLM's embedding space through lightweight linear layers, so every…
The smallest Gemma 4 model, built for efficient on-device execution on phones and laptops, with native audio input.
On-device Gemma 4 model for laptops and mobile devices, with native audio input. 'E' stands for effective parameters.
High-fidelity speech synthesis optimized for quality in structured workflows such as podcasts and audiobooks, with more natural outputs and easier-to-steer prompts.
Fast conversational video generation and editing with native audio, keyframe interpolation and clip extension. Turn text and images into video and refine results through natural language.
Google's first multimodal embedding model, mapping text, images, video, audio, and PDFs into a unified embedding space for semantic search and RAG.


