OpenAI Format

OpenAI Multimodal Responses API

Create responses using the OpenAI Responses API format, with support for multimodal input, tool calling, streaming output, and context compression.

POSThttps://zx1.deepwl.net/v1/responses
Request
curl -X POST https://zx1.deepwl.net/v1/responses \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "gpt-4o",
  "instructions": "You are a concise technical assistant.",
  "input": "Explain in three sentences what scenarios the Responses API is suitable for."
}'
Response
{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1735689600,
  "status": "completed",
  "model": "gpt-4o",
  "output": [
    {
      "type": "message",
      "id": "msg_abc123",
      "status": "completed",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "The Responses API is suitable for multimodal input, tool calling, and tasks that require context continuation. It breaks output into structured items, making it easier for programs to read. Through a unified gateway, you can reuse the same authentication, billing, and channel distribution capabilities.",
          "annotations": []
        }
      ]
    }
  ],
  "usage": {
    "prompt_tokens": 26,
    "completion_tokens": 62,
    "total_tokens": 88
  }
}

The Responses API is a unified response interface for multimodal workflows, tool calling, and context continuation. Compared with Chat Completions, its input, output, and tool-calling structure is better suited for complex task orchestration.

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

Model name, e.g. gpt-4o or a reasoning model.

inputstring | array<object>

Input to the model. Can be a plain text string (equivalent to a text input with the user role), or an array of one or more input items containing different content types such as { "type": "input_text", ... } and { "type": "input_image", ... }.

instructionsstring

A system (or developer) message inserted into the model's context. When used together with previous_response_id, the instructions from a previous response are not carried over to the next response.

reasoningobject

Reasoning configuration, only applicable to reasoning models (o-series and gpt-5 series).

effortstring

Reasoning effort: none, minimal, low, medium, high, xhigh, max. Defaults to medium. Reducing effort can result in faster responses and fewer reasoning tokens. Not all models support every value.

summarystring

A summary of the reasoning performed by the model: auto, concise, or detailed. concise is supported for computer-use-preview models and all reasoning models after gpt-5.

textobject

Text output configuration.

formatobject

Output format. Defaults to { "type": "text" }. Set { "type": "json_schema", "name": ..., "schema": ..., "strict": ... } to enable Structured Outputs, which ensures the model matches your supplied JSON schema. { "type": "json_object" } enables the older JSON mode; json_schema is preferred for models that support it.

verbositystring

Constrains the verbosity of the model's response: low, medium, or high. Defaults to medium.

toolsarray<object>

An array of tools the model may call while generating a response. You can specify which tool to use via tool_choice. Three categories are supported: built-in tools (web_search, file_search, computer_use / computer_use_preview, code_interpreter, image_generation, etc.), MCP tools ({ "type": "mcp", "server_label": ..., ... }), and function tools ({ "type": "function", "name": ..., "description": ..., "parameters": ...(JSON Schema), "strict": ... }). Other upstream-compatible tools can also be forwarded.

tool_choicestring | object

How the model should select which tool(s) to use. String values: none (never call a tool), auto (the model decides, default), required (must call one or more tools). Or pass an object to force a specific tool, e.g. { "type": "function", "name": "my_function" }, { "type": "mcp", "server_label": "..." }, { "type": "allowed_tools", "mode": "auto" | "required", "tools": [...] }, or a hosted tool type such as { "type": "file_search" }, { "type": "web_search_preview" }, { "type": "computer_use_preview" }, { "type": "image_generation" }, { "type": "code_interpreter" }.

parallel_tool_callsboolean默认值: true

Whether to allow the model to run tool calls in parallel. Defaults to true.

max_output_tokensinteger

An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens. Minimum value is 16. Explicitly passing 0 will be preserved and forwarded to supported upstreams.

max_tool_callsinteger

The maximum number of total calls to built-in tools that can be processed in a response, counted across all built-in tool calls. Further tool call attempts by the model beyond this limit will be ignored.

temperaturenumber默认值: 1

Sampling temperature, between 0 and 2. Defaults to 1. Higher values make the output more random. It is generally recommended to alter this or top_p, but not both.

top_pnumber默认值: 1

Nucleus sampling parameter, between 0 and 1. Defaults to 1. For example, 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or temperature, but not both.

top_logprobsinteger

An integer between 0 and 20 specifying the maximum number of most likely tokens to return at each token position, each with an associated log probability.

streamboolean默认值: false

Whether to stream the response with server-sent events (SSE). Defaults to false.

stream_optionsobject

Options for streaming responses. Only set this when stream: true.

include_obfuscationboolean

When true, stream obfuscation is enabled: random characters are added to an obfuscation field on streaming delta events to mitigate certain side-channel attacks. Obfuscation is included by default; set to false to save bandwidth if you trust the network links. Note: in the Responses API, streaming usage is returned on the response.completed event by default and does not depend on stream_options.

backgroundboolean默认值: false

Whether to run the model response in the background. Defaults to false. Background responses can be retrieved later.

storeboolean默认值: true

Whether the upstream stores the generated response for later retrieval via API. Defaults to true. This field is allowed to be forwarded by default and can be disabled by channel settings.

metadataobject

Key-value pairs attached to the response object. Up to 16 pairs; keys up to 64 characters, values up to 512 characters.

includearray<string>

Additional output data to include in the model response. Supported values: web_search_call.results, web_search_call.action.sources, code_interpreter_call.outputs, computer_call_output.output.image_url, file_search_call.results, message.input_image.image_url, message.output_text.logprobs, reasoning.encrypted_content (carries encrypted reasoning content for multi-turn conversations in stateless scenarios, e.g. store: false or zero data retention organizations).

truncationstring默认值: disabled

(Deprecated) The truncation strategy: auto (if the input exceeds the context window, the model truncates by dropping items from the beginning of the conversation) or disabled (fail with a 400 error, default).

previous_response_idstring

The unique ID of the previous response to the model, used to create multi-turn conversations. Cannot be used in conjunction with conversation. Can be used for context continuation when supported by the upstream.

conversationstring | object

The conversation this response belongs to. Can be a conversation ID string or a { "id": "conv_..." } object. Items from this conversation are prepended to the input items for this request, and the input/output items of this response are automatically added to the conversation afterwards. Cannot be used in conjunction with previous_response_id.

promptobject

Reference to a prompt template and its variables.

idstring必填

The unique identifier of the prompt template to use.

versionstring

Optional version of the prompt template.

variablesobject

Optional map of values to substitute into the template's variable placeholders.

prompt_cache_keystring

Used to cache responses for similar requests and optimize cache hit rates. Replaces the deprecated user field.

prompt_cache_optionsobject

Options for prompt caching (supported for gpt-5.6 and later models). Contains ttl (minimum lifetime of cache entries, defaults to 30m) and mode (implicit, the default, automatically creates an implicit cache breakpoint; explicit uses only explicit breakpoints).

safety_identifierstring

A stable identifier used to help detect users of your application that may be violating usage policies. Maximum 64 characters. It is recommended to hash the username or email address to avoid sending identifying information.

service_tierstring默认值: auto

The processing tier used for serving the request: auto (use the Project settings, default), default (standard pricing and performance), flex (Flex Processing), scale, priority / fast (priority processing).

userstring

(Deprecated) A stable identifier for your end-users. Use safety_identifier and prompt_cache_key instead.

moderationobject

Configuration for running moderation on the input and output of this response. Contains model (the moderation model, e.g. omni-moderation-latest, required) and policy (moderation policy for input/output, score or block).

context_managementarray<object>

Context management configuration for this request, controlling how oversized contexts are handled. The exact structure depends on the upstream implementation.

Request Examples

Response Examples

Context Compression

POST /v1/responses/compact is used to compress a long context into a summary suitable for continuing the conversation later. The request structure is similar to /v1/responses, and the commonly used fields are model, input, instructions, and previous_response_id.

curl -X POST https://zx1.deepwl.net/v1/responses/compact \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "gpt-4o",
  "instructions": "Compress this into context that can continue to be used in subsequent conversations, preserving decisions, constraints, and action items.",
  "input": [
    { "role": "user", "content": "First-round requirements..." },
    { "role": "assistant", "content": "First-round proposal..." }
  ]
}'