Anthropic SDK

Anthropic SDK

Claude models can be called with Anthropic's official SDK through the Messages API. Set the base URL to https://api.oxyy.ai without /v1 — the SDK adds /v1/messages itself.

POSThttps://api.oxyy.ai/v1/messages
POSThttps://api.oxyy.ai/v1/messages/count_tokensfree; does not spend your rate limit

Authenticate with your Oxyy API key in the x-api-key header, which is what the SDK's api_key option sends, or as Authorization: Bearer. As on Anthropic's own API, max_tokens is required.

Claude models only. Any other model sent to /v1/messages is refused with a 400, in Anthropic's own error shape so your SDK raises the exception it expects. Use the OpenAI SDK for every other model.

Code examples

# pip install anthropic
import anthropic, os

client = anthropic.Anthropic(
    api_key=os.environ["OXYY_API_KEY"],
    base_url="https://api.oxyy.ai",  # no /v1: the SDK adds it
)

message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system="Be concise.",
    messages=[{"role": "user", "content": "Hello!"}],
)

print(message.content[0].text)
// npm install @anthropic-ai/sdk
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic({
  apiKey: process.env.OXYY_API_KEY,
  baseURL: 'https://api.oxyy.ai', // no /v1: the SDK adds it
});

const message = await client.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 1024,
  messages: [{ role: 'user', content: 'Hello!' }],
});

console.log(message.content[0].text);
curl https://api.oxyy.ai/v1/messages \
  -H "Content-Type: application/json" \
  -H "x-api-key: $OXYY_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
# cache_control survives translation, per block — a long system
# prompt is written once and read at the cache rate afterwards.
message = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": long_style_guide,
        "cache_control": {"type": "ephemeral"},
    }],
    thinking={"type": "enabled", "budget_tokens": 2048},
    messages=[{"role": "user", "content": "Review this draft."}],
)

print(message.usage.cache_read_input_tokens)

Supported features

FeatureStatusNotes
system as a string or a block arraySupportedPer-block cache_control is preserved, so a long system prompt is cached rather than re-billed
max_tokensSupportedRequired, exactly as on Anthropic's own API
temperature, top_p, top_k, stop_sequencesSupported
streamSupportedNative Anthropic events: message_start, content_block_delta, message_delta, message_stop
thinkingSupportedenabled, adaptive and disabled are accepted. On the Claude generations that set thinking depth with an effort level, a budget_tokens is converted to one rather than refused
tools, tool_choiceSupportedauto, any, none and {type:"tool", name}. disable_parallel_tool_use is honoured.
Server tools (web_search, code_execution, …)SupportedPassed to the provider untouched; whether one runs depends on the upstream serving that model
output_config.formatSupportedA json_schema format becomes a structured output; other formats are reported as dropped
metadata, service_tier, betas / anthropic_betaSupportedbetas becomes the anthropic-beta header
container, mcp_servers, context_management, inference_geoSupportedForwarded as sent
POST /v1/messages/count_tokensSupportedAnswered from this gateway's own tokenizer; it does not spend your rate limit
Legacy Text Completions (prompt, max_tokens_to_sample)Not supportedRetired by Anthropic. Refused with a 400 that names the Messages shape to send instead.
Non-Claude modelsNot supportedRefused with a 400. Use the OpenAI SDK for every other model.
Undocumented top-level fieldsNot supportedDropped rather than forwarded, so a validating upstream cannot 400 on a field it has never seen