OpenAI Format

Chat Completions API (Default Streaming)

Start a conversation using the OpenAI Chat Completions-compatible format and return model output via SSE streaming.

POSThttps://zx1.deepwl.net/v1/chat/completions
Request
curl -N -X POST https://zx1.deepwl.net/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "gpt-4o",
  "messages": [
    { "role": "system", "content": "You are a concise technical assistant." },
    { "role": "user", "content": "Explain the role of an API gateway in three points." }
  ],
  "stream": true,
  "stream_options": {
    "include_usage": true
  }
}'
Response
data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{"role":"assistant","content":"API"},"finish_reason":null}]}

data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{"content":" gateway"},"finish_reason":null}]}

data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":26,"completion_tokens":38,"total_tokens":64}}

data: [DONE]

Use a unified conversation format to call upstream models such as OpenAI, Claude, Gemini, DeepSeek, and Qwen. This document uses streaming output as the default, making it suitable for chat, Agent, and long-text generation scenarios where output needs to be displayed as it is generated.

If stream is not provided in this project's route, it is handled as non-streaming. If you want to always receive a streaming response, explicitly pass "stream": true.

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

Model ID, e.g. gpt-4o. You can check the models available to the current API Key via Model List.

messagesarray<object>必填

A list of messages comprising the conversation so far, in chronological order. Different models support different message modalities (text, images, audio, etc.). Common roles are developer (formerly system), system, user, assistant, and tool.

rolestring必填

The role of the message author: developer, system, user, assistant, or tool.

contentstring | array必填

Message content. A string indicates plain text; an array indicates multimodal content parts, supporting text, image_url, input_audio, and file (some upstreams also support video_url). For tool messages, content is the tool result; for assistant messages, content may be null when only tool calls are made.

namestring

An optional name for the participant, helping the model distinguish between speakers with the same role.

tool_callsarray<object>

Only for assistant messages: the tool calls generated by the model. Each item contains id, type (always function), and function (with name and arguments).

tool_call_idstring

Only for tool messages (required): the ID of the tool call this message is responding to.

streamboolean必填

When set to true, the response is text/event-stream, with each chunk pushed as data: and data: [DONE] returned at the end. Officially defaults to false; in this project's route an omitted stream is handled as non-streaming, so pass true explicitly for streaming.

stream_optionsobject

Additional options for streaming output. Only set this when stream: true.

include_usageboolean

Streams an extra chunk containing only usage (with an empty choices array) before data: [DONE], reporting token usage for the entire request. Supported by only some upstream models.

include_obfuscationboolean

Enables stream obfuscation: random characters are added to an obfuscation field on streaming chunks to normalize payload sizes, mitigating certain side-channel attacks. Included by default upstream.

max_completion_tokensinteger

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens. Recommended for reasoning models.

max_tokensinteger

The maximum number of tokens that can be generated. Deprecated upstream in favor of max_completion_tokens, and not compatible with o-series reasoning models.

temperaturenumber

Sampling temperature, between 0 and 2, officially defaults to 1. Higher values make output more random, lower values more focused and deterministic. We generally recommend altering this or top_p but not both.

top_pnumber

Nucleus sampling parameter, between 0 and 1, officially defaults to 1. It is generally not recommended to adjust temperature and top_p significantly at the same time.

ninteger

How many chat completion choices to generate for each input message, between 1 and 128, officially defaults to 1. You are charged for tokens across all choices; keep it at 1 to minimize costs.

stopstring | array<string>

Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence. Not supported with the latest reasoning models such as o3 and o4-mini.

frequency_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the likelihood of repeating the same line verbatim.

presence_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the likelihood of talking about new topics.

logit_biasobject

Modifies the likelihood of specified tokens appearing in the completion. A JSON object mapping token IDs (in the tokenizer) to bias values from -100 to 100; values like -100 / 100 effectively ban / exclusively select the token.

logprobsboolean

Whether to return log probabilities of the output tokens, officially defaults to false. If true, returns the log probability of each output token in the content of message.

top_logprobsinteger

An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with its log probability. logprobs must be set to true when using this parameter.

seedinteger

Random seed (Beta). With the same seed and parameters, the system makes a best effort to sample deterministically, though determinism is not guaranteed; monitor backend changes via the system_fingerprint response field.

response_formatobject

An object specifying the format that the model must output.

typestring

text (default), json_object (the older JSON mode, ensuring valid JSON output), or json_schema (Structured Outputs, ensuring the output matches the supplied JSON schema; preferred for models that support it).

json_schemaobject

Provided when type is json_schema. Contains name and schema (a JSON Schema object), plus optional description and strict.

toolsarray<object>

A list of tools the model may call.

typestring必填

The type of the tool. Function tools are always function; custom tools are also supported.

functionobject必填

The function definition. Contains name (required), description, and parameters (a JSON Schema object), plus optional strict.

tool_choicestring | object

Controls which (if any) tool is called by the model. none means no tool call (default when no tools are present), auto lets the model decide (default when tools are present), required forces one or more tool calls; you can also force a specific function with { "type": "function", "function": { "name": "my_function" } }.

parallel_tool_callsboolean

Whether to enable parallel function calling during tool use, officially defaults to true.

reasoning_effortstring

Constrains effort on reasoning for reasoning models: none, minimal, low, medium, high, xhigh, max; officially defaults to medium. Lower effort yields faster responses and fewer reasoning tokens. Not all reasoning models support every value; effectiveness depends on the model.

verbositystring

Constrains the verbosity of the model's response: low, medium, high; officially defaults to medium. Lower values produce more concise responses. Supported by GPT-5 family models.

modalitiesarray<string>

Output types you would like the model to generate. Defaults to ["text"]; audio models (e.g. gpt-4o-audio-preview) can use ["text", "audio"].

audioobject

Parameters for audio output. Required when modalities includes audio. Contains voice (built-in voices such as alloy, ash, coral, echo, nova, onyx, sage, shimmer, etc.) and format (wav, mp3, flac, opus, pcm16, aac).

predictionobject

Configuration for a Predicted Output, which can greatly improve response times when large parts of the response are known ahead of time (e.g. regenerating a file with minor changes). Shape: { "type": "content", "content": ... }.

storeboolean

Whether to store the output of this request (used upstream for distillation / evals products), officially defaults to false.

metadataobject

Set of up to 16 key-value pairs for storing additional structured information. Keys are strings up to 64 characters; values up to 512 characters.

service_tierstring

The processing tier used for serving the request: auto (default, uses the project setting), default, flex, scale, priority, fast. The service_tier in the response reflects the tier actually used and may differ from the request value. Effectiveness depends on the upstream.

web_search_optionsobject

Web search options (for search-capable models). Supports search_context_size (low / medium / high) and user_location (approximate location).

prompt_cache_keystring

A cache key used to optimize cache hit rates for similar requests; officially replaces the user field.

safety_identifierstring

A stable identifier used to help detect end users who may be violating usage policies, up to 64 characters. Hashing the username or email is recommended to avoid sending identifying information.

userstring

A stable identifier for your end users. Deprecated upstream in favor of safety_identifier and prompt_cache_key.

Request Examples

Response Examples

Response Fields

idstring

The response ID generated for this request.

objectstring

The streaming response is always chat.completion.chunk.

choicesarray<object>

Array of candidate results.

deltaobject

Incremental content. May include role, content, reasoning_content, or tool_calls.

finish_reasonstring

Reason for completion. Common values are stop, length, and tool_calls.

usageobject

Usage statistics. It is guaranteed to appear only when the upstream returns usage data and stream_options.include_usage is enabled.