oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. Google/
  4. Gemma 4 26B A4B Instruct (MoE)

Google: Gemma 4 26B A4B Instruct (MoE)

google/gemma-4-26b-a4b-it
Compare Playground Get API key

Gemma 4's Mixture-of-Experts model. 25.2B total parameters but only 3.8B active per token, so it runs almost as fast as a 4B model while scoring close to the dense 31B.

ModalitiesTextImageVideoText
In / out price$0.052 / $0.160 per 1M
Context262K
ReleasedApr 2, 2026
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
Google$0.052$0.160$09.93s55 tps

Capabilities (13)

StreamingTool callingStructured outputThinkingReasoningSystem promptVisionVideo InputPDF inputMultilingualInterleaved Image TextPrompt cachingFine Tunable

Supported parameters (7)

thinking_leveltemperaturetop_ptop_ktoolsresponse_formatreasoning_effort

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
$0.052
/M tokens · listed $0.130
Effective output price
$0.160
/M tokens · listed $0.400
RateListedChargedUnit
Input$0.130$0.052per 1M tokens
Output$0.400$0.160per 1M tokens
Cache read$0$0per 1M tokens
Cache write$0$0per 1M tokens
Cache write (1h)$0$0per 1M tokens
Cached audio read$0$0per 1M tokens
Cache storage$0$0per 1M token-hours
Reasoning$0$0per 1M tokens
UTC

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
55tok/s
P50, streaming requests
Latency
9.93s
P50, end to end
UTC

Throughput

P99198 tok/s
P9593 tok/s
P5058 tok/s

Latency

P5011.1s
P7518.7s
P9554.8s
P99172.2s

Time to first token

P501353ms
P956395ms

Cache hit rate

4.90%
Share of input tokens served from the prompt cache over the window.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
93.10%
over 21,884 requests
Days with traffic
31
of the last 30

Days are UTC days. Days with no requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

  1. 1.Agnai.Chat 4.2K tokens

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt51.7M
Reasoning1.9M
Completion19.6M

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length.

Frequently asked questions

What is Gemma 4 26B A4B Instruct (MoE)?
Gemma 4's Mixture-of-Experts model. 25.2B total parameters but only 3.8B active per token, so it runs almost as fast as a 4B model while scoring close to the dense 31B.
How much does Gemma 4 26B A4B Instruct (MoE) cost?
Gemma 4 26B A4B Instruct (MoE) costs $0.130 per million input tokens and $0.400 per million output tokens. Cached prompt tokens are billed at $0 per million. You pay per request, with no subscription.
What is the context length of Gemma 4 26B A4B Instruct (MoE)?
Gemma 4 26B A4B Instruct (MoE) has a context window of 262,144 tokens.
Does Gemma 4 26B A4B Instruct (MoE) support tool calling and structured outputs?
Yes — Gemma 4 26B A4B Instruct (MoE) supports both tool calling and structured outputs.
What inputs and outputs does Gemma 4 26B A4B Instruct (MoE) support?
Gemma 4 26B A4B Instruct (MoE) accepts text, image, video as input and returns text.
What other models does Google have?
Google also offers Gemini 2.5 Flash, Gemini 2.5 Flash-Lite, Gemini 2.5 Flash TTS, Gemini 2.5 Pro through Oxyy.
When was Gemma 4 26B A4B Instruct (MoE) released?
Gemma 4 26B A4B Instruct (MoE) was released on Apr 2, 2026.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from Google

Gemini 2.5 FlashGoogle's first hybrid reasoning model with configurable thinking budgets; best price-performance for low-latency, high-…text · 1.05M context · $0.300 / $2.50Gemini 2.5 Flash-LiteSmallest and most cost-effective multimodal model in the 2.5 family, built for at-scale usage.text · 1.05M context · $0.100 / $0.400Gemini 2.5 Flash TTSFast and controllable text-to-speech for low-latency, cost-efficient applications and real-time assistants, with fine c…audio tts · 8K context · $0.500 / —Gemini 2.5 ProMost advanced model of the 2.5 family, with deep reasoning and coding capability for complex tasks.text · 1.05M context · $1.25 / $10.00Gemini 2.5 Pro TTSHigh-fidelity speech synthesis optimized for quality in structured workflows such as podcasts and audiobooks, with more…audio tts · 8K context · $1.00 / —Nano Banana 2High-efficiency production-scale image generation and editing, balancing speed with 4K generation, world knowledge and…image generation · — context · $0.500 / $3.00
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout