oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. Qwen (Alibaba)/
  4. Qwen3 30B A3B (FP8)

Qwen (Alibaba): Qwen3 30B A3B (FP8)

qwen/qwen3-30b-a3b-fp8
Compare Playground Get API key
FP8-quantized checkpoint of Qwen3-30B-A3B, a Mixture-of-Experts model with hybrid thinking and non-thinking modes.
ModalitiesTextText
In / out price— / — per 1M
Context33K
ReleasedApr 28, 2025
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
Qwen (Alibaba)———1.33s175 tps

Capabilities (6)

StreamingTool callingStructured outputThinkingReasoningMultilingual

Supported parameters (3)

enable_thinkingtoolsresponse_format

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
—
/M tokens
Effective output price
—
/M tokens
RateListedChargedUnit
UTC
Price history starts building from the first daily snapshot — there is not enough of it yet to draw a trend.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
175tok/s
P50, streaming requests
Latency
1.33s
P50, end to end
UTC

Throughput

P99191 tok/s
P95190 tok/s
P50161 tok/s

Latency

P501374ms
P751507ms
P952000ms
P992202ms

Time to first token

P50588ms
P95930ms

Cache hit rate

0.00%
Share of input tokens served from the prompt cache over the window.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
83.30%
over 390 requests
Days with traffic
29
of the last 30

Days are UTC days. Days with no requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

No apps have identified themselves for this model yet. Clients can opt in by sending an X-Title header with their request.

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt12.5K
Completion45.1K

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is Qwen3 30B A3B (FP8)?
FP8-quantized checkpoint of Qwen3-30B-A3B, a Mixture-of-Experts model with hybrid thinking and non-thinking modes.
How much does Qwen3 30B A3B (FP8) cost?
Pricing for Qwen3 30B A3B (FP8) is listed on this page under Pricing. You pay per request, with no subscription.
What is the context length of Qwen3 30B A3B (FP8)?
Qwen3 30B A3B (FP8) has a context window of 32,768 tokens.
Does Qwen3 30B A3B (FP8) support tool calling and structured outputs?
Yes — Qwen3 30B A3B (FP8) supports both tool calling and structured outputs.
What inputs and outputs does Qwen3 30B A3B (FP8) support?
Qwen3 30B A3B (FP8) accepts text as input and returns text.
When was Qwen3 30B A3B (FP8) released?
Qwen3 30B A3B (FP8) was released on Apr 28, 2025.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout