oxyy.ai
HomeModelsProvidersPricingBlogDocs
Sign inGet API key
  1. Home/
  2. Providers/
  3. Z.ai/
  4. GLM-4-32B-0414-128K

Z.ai: GLM-4-32B-0414-128K

zai/glm-4-32b-0414-128k
Compare Playground Get API key
32-billion-parameter GLM-4 model with a 128K context window, flat-priced on input and output.
ModalitiesTextText
In / out price$0.040 / $0.040 per 1M
Context128K
Released—
ProvidersPricingPerformanceAvailabilityAppsActivityFAQExplore

Providers

Where this model runs, and what each route costs. Requests are sent to a healthy provider automatically; if one errors, the gateway retries against another serving the same model.

ProviderInput /MOutput /MCache read /MLatencyThroughput
Z.ai$0.040$0.040———

Capabilities (2)

StreamingTool calling

Supported parameters (1)

tools

Any other parameter you send is ignored rather than rejected.

Pricing

What this model costs to run, next to the rate it is posted at. Caching and discounts mean the price actually paid is often below the listed one.

Effective input price
$0.040
/M tokens · listed $0.100
Effective output price
$0.040
/M tokens · listed $0.100
RateListedChargedUnit
Input$0.100$0.040per 1M tokens
Output$0.100$0.040per 1M tokens
UTC

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better). Throughput and TTFT are measured on streaming requests, which are the only ones with a first-token moment to time.

Throughput
—tok/s
P50, streaming requests
Latency
—s
P50, end to end
UTC
No requests to this model in the selected window, so there is nothing to measure yet.

Availability

The share of requests to this model that completed successfully. When an upstream provider errors, the gateway retries against another provider serving the same model, so a single provider incident does not necessarily show up here.

Success rate (30d)
0.00%
over 83 requests
Days with traffic
2
of the last 30

Days are UTC days. Days with no requests are omitted rather than drawn at 100% — no traffic is not evidence of availability.

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

  1. 1.Oxyy Tester112 tokens

Activity

Token volume and request traffic to this model over time. Daily totals, UTC.

Prompt112
Completion0

Prompt tokens measure input size. Reasoning tokens show internal thinking before a response. Completion tokens reflect total output length. Reasoning is not priced separately for this model, so it is not itemised here.

Frequently asked questions

What is GLM-4-32B-0414-128K?
32-billion-parameter GLM-4 model with a 128K context window, flat-priced on input and output.
How much does GLM-4-32B-0414-128K cost?
GLM-4-32B-0414-128K costs $0.100 per million input tokens and $0.100 per million output tokens. You pay per request, with no subscription.
What is the context length of GLM-4-32B-0414-128K?
GLM-4-32B-0414-128K has a context window of 128,000 tokens.
Does GLM-4-32B-0414-128K support tool calling and structured outputs?
GLM-4-32B-0414-128K supports tool calling. Structured outputs are not listed among its capabilities.
What inputs and outputs does GLM-4-32B-0414-128K support?
GLM-4-32B-0414-128K accepts text as input and returns text.
What other models does Z.ai have?
Z.ai also offers CogVideoX-3, CogView-4, GLM-4.5, GLM-4.5-Air through Oxyy.
When was GLM-4-32B-0414-128K released?
No release date is published for GLM-4-32B-0414-128K.

Explore more models

Browse the full catalog ModelsCompare pricing Pricing

More models from Z.ai

CogVideoX-3Z.ai's video generation model.text to video · — contextCogView-4Z.ai's CogView image generation model.image generation · — contextGLM-4.5GLM-4.5 base text model. The release that introduced interleaved reasoning to the GLM line.text · — context · $0.600 / $2.20GLM-4.5-AirLightweight, low-cost variant of GLM-4.5.text · — context · $0.200 / $1.10GLM-4.5-AirXHigh-throughput variant of GLM-4.5-Air.text · — context · $1.10 / $4.50GLM-4.5-FlashFree tier text model from the GLM-4.5 generation.text · — context · $0 / $0
oxyy.ai

One endpoint for every major model. The OpenAI, Anthropic and Gemini SDKs work unchanged.

Product

  • Models
  • Providers
  • Pricing
  • Status
  • Startup program

Company

  • About
  • Blog
  • Contact
  • Terms of Service
  • Privacy Policy
  • Refund Policy
  • Cookie Policy

Developer

  • Documentation
  • Quickstart
  • SDKs
  • Model catalog
  • AI providers

Connect

  • Telegram
  • Discord
  • WhatsApp
© 2026 Oxyy.ai. All rights reserved.StatusAbout