OpenAI Multimodal Responses API
Create responses using the OpenAI Responses API format, with support for multimodal input, tool calling, streaming output, and context compression.
https://zx1.deepwl.net/v1/responsescurl -X POST https://zx1.deepwl.net/v1/responses \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"model": "gpt-4o",
"instructions": "You are a concise technical assistant.",
"input": "Explain in three sentences what scenarios the Responses API is suitable for."
}'{
"id": "resp_abc123",
"object": "response",
"created_at": 1735689600,
"status": "completed",
"model": "gpt-4o",
"output": [
{
"type": "message",
"id": "msg_abc123",
"status": "completed",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "The Responses API is suitable for multimodal input, tool calling, and tasks that require context continuation. It breaks output into structured items, making it easier for programs to read. Through a unified gateway, you can reuse the same authentication, billing, and channel distribution capabilities.",
"annotations": []
}
]
}
],
"usage": {
"prompt_tokens": 26,
"completion_tokens": 62,
"total_tokens": 88
}
}The Responses API is a unified response interface for multimodal workflows, tool calling, and context continuation. Compared with Chat Completions, its input, output, and tool-calling structure is better suited for complex task orchestration.
Request Headers
Authorizationstring必填Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
Content-Typestring必填Request content type, must be application/json.
Request Body
modelstring必填Model name, e.g. gpt-4o or a reasoning model.
inputstring | array<object>Input to the model. Can be a plain text string (equivalent to a text input with the user role), or an array of one or more input items containing different content types such as { "type": "input_text", ... } and { "type": "input_image", ... }.
instructionsstringA system (or developer) message inserted into the model's context. When used together with previous_response_id, the instructions from a previous response are not carried over to the next response.
reasoningobjectReasoning configuration, only applicable to reasoning models (o-series and gpt-5 series).
effortstringReasoning effort: none, minimal, low, medium, high, xhigh, max. Defaults to medium. Reducing effort can result in faster responses and fewer reasoning tokens. Not all models support every value.
summarystringA summary of the reasoning performed by the model: auto, concise, or detailed. concise is supported for computer-use-preview models and all reasoning models after gpt-5.
textobjectText output configuration.
formatobjectOutput format. Defaults to { "type": "text" }. Set { "type": "json_schema", "name": ..., "schema": ..., "strict": ... } to enable Structured Outputs, which ensures the model matches your supplied JSON schema. { "type": "json_object" } enables the older JSON mode; json_schema is preferred for models that support it.
verbositystringConstrains the verbosity of the model's response: low, medium, or high. Defaults to medium.
toolsarray<object>An array of tools the model may call while generating a response. You can specify which tool to use via tool_choice. Three categories are supported: built-in tools (web_search, file_search, computer_use / computer_use_preview, code_interpreter, image_generation, etc.), MCP tools ({ "type": "mcp", "server_label": ..., ... }), and function tools ({ "type": "function", "name": ..., "description": ..., "parameters": ...(JSON Schema), "strict": ... }). Other upstream-compatible tools can also be forwarded.
tool_choicestring | objectHow the model should select which tool(s) to use. String values: none (never call a tool), auto (the model decides, default), required (must call one or more tools). Or pass an object to force a specific tool, e.g. { "type": "function", "name": "my_function" }, { "type": "mcp", "server_label": "..." }, { "type": "allowed_tools", "mode": "auto" | "required", "tools": [...] }, or a hosted tool type such as { "type": "file_search" }, { "type": "web_search_preview" }, { "type": "computer_use_preview" }, { "type": "image_generation" }, { "type": "code_interpreter" }.
parallel_tool_callsboolean默认值: trueWhether to allow the model to run tool calls in parallel. Defaults to true.
max_output_tokensintegerAn upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens. Minimum value is 16. Explicitly passing 0 will be preserved and forwarded to supported upstreams.
max_tool_callsintegerThe maximum number of total calls to built-in tools that can be processed in a response, counted across all built-in tool calls. Further tool call attempts by the model beyond this limit will be ignored.
temperaturenumber默认值: 1Sampling temperature, between 0 and 2. Defaults to 1. Higher values make the output more random. It is generally recommended to alter this or top_p, but not both.
top_pnumber默认值: 1Nucleus sampling parameter, between 0 and 1. Defaults to 1. For example, 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or temperature, but not both.
top_logprobsintegerAn integer between 0 and 20 specifying the maximum number of most likely tokens to return at each token position, each with an associated log probability.
streamboolean默认值: falseWhether to stream the response with server-sent events (SSE). Defaults to false.
stream_optionsobjectOptions for streaming responses. Only set this when stream: true.
include_obfuscationbooleanWhen true, stream obfuscation is enabled: random characters are added to an obfuscation field on streaming delta events to mitigate certain side-channel attacks. Obfuscation is included by default; set to false to save bandwidth if you trust the network links. Note: in the Responses API, streaming usage is returned on the response.completed event by default and does not depend on stream_options.
backgroundboolean默认值: falseWhether to run the model response in the background. Defaults to false. Background responses can be retrieved later.
storeboolean默认值: trueWhether the upstream stores the generated response for later retrieval via API. Defaults to true. This field is allowed to be forwarded by default and can be disabled by channel settings.
metadataobjectKey-value pairs attached to the response object. Up to 16 pairs; keys up to 64 characters, values up to 512 characters.
includearray<string>Additional output data to include in the model response. Supported values: web_search_call.results, web_search_call.action.sources, code_interpreter_call.outputs, computer_call_output.output.image_url, file_search_call.results, message.input_image.image_url, message.output_text.logprobs, reasoning.encrypted_content (carries encrypted reasoning content for multi-turn conversations in stateless scenarios, e.g. store: false or zero data retention organizations).
truncationstring默认值: disabled(Deprecated) The truncation strategy: auto (if the input exceeds the context window, the model truncates by dropping items from the beginning of the conversation) or disabled (fail with a 400 error, default).
previous_response_idstringThe unique ID of the previous response to the model, used to create multi-turn conversations. Cannot be used in conjunction with conversation. Can be used for context continuation when supported by the upstream.
conversationstring | objectThe conversation this response belongs to. Can be a conversation ID string or a { "id": "conv_..." } object. Items from this conversation are prepended to the input items for this request, and the input/output items of this response are automatically added to the conversation afterwards. Cannot be used in conjunction with previous_response_id.
promptobjectReference to a prompt template and its variables.
idstring必填The unique identifier of the prompt template to use.
versionstringOptional version of the prompt template.
variablesobjectOptional map of values to substitute into the template's variable placeholders.
prompt_cache_keystringUsed to cache responses for similar requests and optimize cache hit rates. Replaces the deprecated user field.
prompt_cache_optionsobjectOptions for prompt caching (supported for gpt-5.6 and later models). Contains ttl (minimum lifetime of cache entries, defaults to 30m) and mode (implicit, the default, automatically creates an implicit cache breakpoint; explicit uses only explicit breakpoints).
safety_identifierstringA stable identifier used to help detect users of your application that may be violating usage policies. Maximum 64 characters. It is recommended to hash the username or email address to avoid sending identifying information.
service_tierstring默认值: autoThe processing tier used for serving the request: auto (use the Project settings, default), default (standard pricing and performance), flex (Flex Processing), scale, priority / fast (priority processing).
userstring(Deprecated) A stable identifier for your end-users. Use safety_identifier and prompt_cache_key instead.
moderationobjectConfiguration for running moderation on the input and output of this response. Contains model (the moderation model, e.g. omni-moderation-latest, required) and policy (moderation policy for input/output, score or block).
context_managementarray<object>Context management configuration for this request, controlling how oversized contexts are handled. The exact structure depends on the upstream implementation.
Request Examples
Response Examples
Context Compression
POST /v1/responses/compact is used to compress a long context into a summary suitable for continuing the conversation later. The request structure is similar to /v1/responses, and the commonly used fields are model, input, instructions, and previous_response_id.
curl -X POST https://zx1.deepwl.net/v1/responses/compact \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"model": "gpt-4o",
"instructions": "Compress this into context that can continue to be used in subsequent conversations, preserving decisions, constraints, and action items.",
"input": [
{ "role": "user", "content": "First-round requirements..." },
{ "role": "assistant", "content": "First-round proposal..." }
]
}'