Claude Messages API
Call Claude models using Anthropic Messages native format.
https://zx1.deepwl.net/v1/messagescurl -X POST https://zx1.deepwl.net/v1/messages \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"system": "You are a rigorous technical advisor.",
"messages": [
{
"role": "user",
"content": "Explain why an API gateway needs rate limiting."
}
]
}'{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"model": "claude-sonnet-4-20250514",
"content": [
{
"type": "text",
"text": "API gateway rate limiting can protect upstream services, prevent sudden traffic from exhausting resources, and provide stable service quality for different users."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 32,
"output_tokens": 45
}
}The Claude Messages API preserves Anthropic's native request structure and is suitable for migrating existing Claude SDK-based services or services using native prompt structures. Requests will be routed into the Claude relay format and dispatched to the corresponding upstream based on the channel.
Request Headers
Authorizationstring必填Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
x-api-keystringClaude native authentication header, can be used instead of Authorization, e.g. YOUR_API_KEY.
anthropic-versionstringClaude native version header, used together with x-api-key, e.g. 2023-06-01.
Content-Typestring必填Request content type, must be application/json.
Request Body
modelstring必填Claude model name, e.g. claude-sonnet-4-20250514.
max_tokensinteger必填Maximum number of output tokens. Required by the official API; generation stops when this limit is reached (stop_reason is max_tokens).
messagesarray<object>必填Array of conversation messages. Roles must alternate between user and assistant.
rolestring必填Message role: user or assistant.
contentstring | array<object>必填Message content. Can be a plain string or an array of content blocks. Common block types: text, image (with a base64 or url source), document, tool_use, tool_result, thinking.
systemstring | array<object>System prompt. Claude does not use system role messages; instead, system instructions are passed through this field. When passed as an array, elements are text content blocks that can be combined with cache_control for prompt caching.
streambooleanWhether to enable SSE streaming output. Defaults to false.
temperaturenumberSampling temperature, range 0 to 1, defaults to 1.0. Lower values produce more deterministic output. Anthropic recommends adjusting either temperature or top_p, but not both.
top_pnumberNucleus sampling parameter. Avoid adjusting it together with temperature.
top_kintegerTop-K sampling parameter. Only samples from the top K most probable tokens.
stop_sequencesarray<string>Custom stop sequences. The model stops after generating any of these sequences; the returned text does not include the sequence itself.
toolsarray<object>List of tool definitions.
namestring必填Tool name.
descriptionstringDescription of what the tool does, helping the model decide when to call it.
input_schemaobject必填JSON Schema for the tool parameters, typically { "type": "object", "properties": {...}, "required": [...] }.
tool_choiceobjectControls the tool selection strategy.
typestring必填auto (model decides whether to call, default), any (must call some tool), tool (force a specific tool), none (disable tool use).
namestringWhen type is tool, the name of the tool to force.
disable_parallel_tool_usebooleanSet to true to prevent the model from calling multiple tools in parallel within one response.
thinkingobjectExtended thinking configuration. Only available on models with thinking capability.
typestring必填enabled to turn thinking on, disabled to turn it off.
budget_tokensintegerToken budget for thinking. Must be less than max_tokens.
metadataobjectRequest metadata.
user_idstringA stable identifier for the end user (preferably an irreversible ID or hash), used for abuse detection.
service_tierstringService capacity tier: auto (default; uses priority capacity when available, falling back to standard), standard_only (standard capacity only).
output_configobjectStructured output configuration. Pass { "type": "json_schema", "schema": {...} } in format to force the response to conform to the given JSON Schema.
context_managementobjectContext management configuration. The edits array defines context editing strategies (e.g. automatically clearing old tool calls and results) to control the context window in long conversations.
mcp_serversarray<object>List of MCP connector servers. Each item includes type (always url), name, url, and optionally authorization_token and tool_configuration (with enabled, allowed_tools).
containerstringContainer ID. Pass a container identifier returned by a previous response to reuse the code execution container across requests.
inference_geostringInference geography restriction, e.g. "us" to run inference only in the United States.
Request Examples
Response Examples
Response Fields
idstringThe message ID.
typestringObject type, always message.
rolestringMessage role, always assistant.
modelstringThe model that generated the response.
contentarray<object>Array of content blocks.
typestringContent block type, e.g. text.
textstringThe text content generated by the model.
stop_reasonstringStop reason, e.g. end_turn.
usageobjectToken usage statistics.
input_tokensintegerNumber of input tokens.
output_tokensintegerNumber of output tokens.