oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. Z.ai/
  4. GLM-5.2

Z.ai: GLM-5.2

zai/glm-5.2
Compare Playground Get API key
Long-context reasoning and coding model. The GLM release that expanded the context window from 200K to 1M tokens.
ModalitiesTextText
In / out price$0.560 / $1.76 per 1M
Context1.05M
ReleasedJun 16, 2026
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
Z.ai$0.560$1.76$0.26019.59s70 tps

Capabilities (9)

StreamingTool callingTool StreamingStructured outputThinkingReasoningSystem promptPrompt cachingWeb search

Supported parameters (2)

toolsresponse_format

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
$0.560
/M tokens · listed $1.40
Effective output price
$1.76
/M tokens · listed $4.40
RateListedChargedUnit
Input$1.40$0.560per 1M tokens
Output$4.40$1.76per 1M tokens
Cache read$0.260$0.260per 1M tokens
UTC

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
70tok/s
P50, streaming requests
Latency
19.59s
P50, end to end
UTC

Throughput

P99674 tok/s
P95114 tok/s
P5068 tok/s

Latency

P5019.6s
P7537.5s
P9582.8s
P99141.8s

Time to first token

P503882ms
P9516.8s

Cache hit rate

8.54%
Share of input tokens served from the prompt cache over the window.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
—
no requests in the window
Days with traffic
0
of the last 30
This model has not served any requests in the last 30 days, so there is no availability to report.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

  1. 1.Z Code@electron 11.8M tokens
  2. 2.Agnai.Chat 6.9K tokens
  3. 3.Oxyy Tester56 tokens

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt12.4B
Reasoning7.1M
Completion536.4M

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length.

Frequently asked questions

What is GLM-5.2?
Long-context reasoning and coding model. The GLM release that expanded the context window from 200K to 1M tokens.
How much does GLM-5.2 cost?
GLM-5.2 costs $1.40 per million input tokens and $4.40 per million output tokens. Cached prompt tokens are billed at $0.260 per million. You pay per request, with no subscription.
What is the context length of GLM-5.2?
GLM-5.2 has a context window of 1,048,576 tokens, and can return up to 128,000 tokens in one response.
Does GLM-5.2 support tool calling and structured outputs?
Yes — GLM-5.2 supports both tool calling and structured outputs.
What inputs and outputs does GLM-5.2 support?
GLM-5.2 accepts text as input and returns text.
What other models does Z.ai have?
Z.ai also offers CogVideoX-3, CogView-4, GLM-4-32B-0414-128K, GLM-4.5 through Oxyy.
When was GLM-5.2 released?
GLM-5.2 was released on Jun 16, 2026.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from Z.ai

CogVideoX-3Z.ai's video generation model.text to video · — contextCogView-4Z.ai's CogView image generation model.image generation · — contextGLM-4-32B-0414-128K32-billion-parameter GLM-4 model with a 128K context window, flat-priced on input and output.text · 128K context · $0.100 / $0.100GLM-4.5GLM-4.5 base text model. The release that introduced interleaved reasoning to the GLM line.text · — context · $0.600 / $2.20GLM-4.5-AirLightweight, low-cost variant of GLM-4.5.text · — context · $0.200 / $1.10GLM-4.5-AirXHigh-throughput variant of GLM-4.5-Air.text · — context · $1.10 / $4.50
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout