Anthropic SDK
Anthropic SDK
Claude models can be called with Anthropic's official SDK through the Messages API. Set the base URL to https://api.oxyy.ai without /v1 — the SDK adds /v1/messages itself.
POSThttps://api.oxyy.ai/v1/messages
POSThttps://api.oxyy.ai/v1/messages/count_tokensfree; does not spend your rate limit
Authenticate with your Oxyy API key in the x-api-key header, which is what the SDK's api_key option sends, or as Authorization: Bearer. As on Anthropic's own API, max_tokens is required.
Claude models only. Any other model sent to
/v1/messages is refused with a 400, in Anthropic's own error shape so your SDK raises the exception it expects. Use the OpenAI SDK for every other model.Code examples
# pip install anthropic import anthropic, os client = anthropic.Anthropic( api_key=os.environ["OXYY_API_KEY"], base_url="https://api.oxyy.ai", # no /v1: the SDK adds it ) message = client.messages.create( model="claude-sonnet-5", max_tokens=1024, system="Be concise.", messages=[{"role": "user", "content": "Hello!"}], ) print(message.content[0].text)
// npm install @anthropic-ai/sdk import Anthropic from '@anthropic-ai/sdk'; const client = new Anthropic({ apiKey: process.env.OXYY_API_KEY, baseURL: 'https://api.oxyy.ai', // no /v1: the SDK adds it }); const message = await client.messages.create({ model: 'claude-sonnet-5', max_tokens: 1024, messages: [{ role: 'user', content: 'Hello!' }], }); console.log(message.content[0].text);
curl https://api.oxyy.ai/v1/messages \ -H "Content-Type: application/json" \ -H "x-api-key: $OXYY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -d '{ "model": "claude-sonnet-5", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello!"}] }'
# cache_control survives translation, per block — a long system # prompt is written once and read at the cache rate afterwards. message = client.messages.create( model="claude-sonnet-5", max_tokens=1024, system=[{ "type": "text", "text": long_style_guide, "cache_control": {"type": "ephemeral"}, }], thinking={"type": "enabled", "budget_tokens": 2048}, messages=[{"role": "user", "content": "Review this draft."}], ) print(message.usage.cache_read_input_tokens)
Supported features
| Feature | Status | Notes |
|---|---|---|
system as a string or a block array | Supported | Per-block cache_control is preserved, so a long system prompt is cached rather than re-billed |
max_tokens | Supported | Required, exactly as on Anthropic's own API |
temperature, top_p, top_k, stop_sequences | Supported | |
stream | Supported | Native Anthropic events: message_start, content_block_delta, message_delta, message_stop |
thinking | Supported | enabled, adaptive and disabled are accepted. On the Claude generations that set thinking depth with an effort level, a budget_tokens is converted to one rather than refused |
tools, tool_choice | Supported | auto, any, none and {type:"tool", name}. disable_parallel_tool_use is honoured. |
Server tools (web_search, code_execution, …) | Supported | Passed to the provider untouched; whether one runs depends on the upstream serving that model |
output_config.format | Supported | A json_schema format becomes a structured output; other formats are reported as dropped |
metadata, service_tier, betas / anthropic_beta | Supported | betas becomes the anthropic-beta header |
container, mcp_servers, context_management, inference_geo | Supported | Forwarded as sent |
POST /v1/messages/count_tokens | Supported | Answered from this gateway's own tokenizer; it does not spend your rate limit |
Legacy Text Completions (prompt, max_tokens_to_sample) | Not supported | Retired by Anthropic. Refused with a 400 that names the Messages shape to send instead. |
| Non-Claude models | Not supported | Refused with a 400. Use the OpenAI SDK for every other model. |
| Undocumented top-level fields | Not supported | Dropped rather than forwarded, so a validating upstream cannot 400 on a field it has never seen |
