Chat Completions API (Default Streaming)
Start a conversation using the OpenAI Chat Completions-compatible format and return model output via SSE streaming.
https://zx1.deepwl.net/v1/chat/completionscurl -N -X POST https://zx1.deepwl.net/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"model": "gpt-4o",
"messages": [
{ "role": "system", "content": "You are a concise technical assistant." },
{ "role": "user", "content": "Explain the role of an API gateway in three points." }
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{"role":"assistant","content":"API"},"finish_reason":null}]}
data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{"content":" gateway"},"finish_reason":null}]}
data: {"id":"chatcmpl_abc123","object":"chat.completion.chunk","created":1735689600,"model":"gpt-4o","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":26,"completion_tokens":38,"total_tokens":64}}
data: [DONE]Use a unified conversation format to call upstream models such as OpenAI, Claude, Gemini, DeepSeek, and Qwen. This document uses streaming output as the default, making it suitable for chat, Agent, and long-text generation scenarios where output needs to be displayed as it is generated.
If stream is not provided in this project's route, it is handled as non-streaming. If you want to always receive a streaming response, explicitly pass "stream": true.
Request Headers
Authorizationstring必填Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
Content-Typestring必填Request content type, must be application/json.
Request Body
modelstring必填Model ID, e.g. gpt-4o. You can check the models available to the current API Key via Model List.
messagesarray<object>必填A list of messages comprising the conversation so far, in chronological order. Different models support different message modalities (text, images, audio, etc.). Common roles are developer (formerly system), system, user, assistant, and tool.
rolestring必填The role of the message author: developer, system, user, assistant, or tool.
contentstring | array必填Message content. A string indicates plain text; an array indicates multimodal content parts, supporting text, image_url, input_audio, and file (some upstreams also support video_url). For tool messages, content is the tool result; for assistant messages, content may be null when only tool calls are made.
namestringAn optional name for the participant, helping the model distinguish between speakers with the same role.
tool_callsarray<object>Only for assistant messages: the tool calls generated by the model. Each item contains id, type (always function), and function (with name and arguments).
tool_call_idstringOnly for tool messages (required): the ID of the tool call this message is responding to.
streamboolean必填When set to true, the response is text/event-stream, with each chunk pushed as data: and data: [DONE] returned at the end. Officially defaults to false; in this project's route an omitted stream is handled as non-streaming, so pass true explicitly for streaming.
stream_optionsobjectAdditional options for streaming output. Only set this when stream: true.
include_usagebooleanStreams an extra chunk containing only usage (with an empty choices array) before data: [DONE], reporting token usage for the entire request. Supported by only some upstream models.
include_obfuscationbooleanEnables stream obfuscation: random characters are added to an obfuscation field on streaming chunks to normalize payload sizes, mitigating certain side-channel attacks. Included by default upstream.
max_completion_tokensintegerAn upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens. Recommended for reasoning models.
max_tokensintegerThe maximum number of tokens that can be generated. Deprecated upstream in favor of max_completion_tokens, and not compatible with o-series reasoning models.
temperaturenumberSampling temperature, between 0 and 2, officially defaults to 1. Higher values make output more random, lower values more focused and deterministic. We generally recommend altering this or top_p but not both.
top_pnumberNucleus sampling parameter, between 0 and 1, officially defaults to 1. It is generally not recommended to adjust temperature and top_p significantly at the same time.
nintegerHow many chat completion choices to generate for each input message, between 1 and 128, officially defaults to 1. You are charged for tokens across all choices; keep it at 1 to minimize costs.
stopstring | array<string>Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence. Not supported with the latest reasoning models such as o3 and o4-mini.
frequency_penaltynumberNumber between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the likelihood of repeating the same line verbatim.
presence_penaltynumberNumber between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the likelihood of talking about new topics.
logit_biasobjectModifies the likelihood of specified tokens appearing in the completion. A JSON object mapping token IDs (in the tokenizer) to bias values from -100 to 100; values like -100 / 100 effectively ban / exclusively select the token.
logprobsbooleanWhether to return log probabilities of the output tokens, officially defaults to false. If true, returns the log probability of each output token in the content of message.
top_logprobsintegerAn integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with its log probability. logprobs must be set to true when using this parameter.
seedintegerRandom seed (Beta). With the same seed and parameters, the system makes a best effort to sample deterministically, though determinism is not guaranteed; monitor backend changes via the system_fingerprint response field.
response_formatobjectAn object specifying the format that the model must output.
typestringtext (default), json_object (the older JSON mode, ensuring valid JSON output), or json_schema (Structured Outputs, ensuring the output matches the supplied JSON schema; preferred for models that support it).
json_schemaobjectProvided when type is json_schema. Contains name and schema (a JSON Schema object), plus optional description and strict.
toolsarray<object>A list of tools the model may call.
typestring必填The type of the tool. Function tools are always function; custom tools are also supported.
functionobject必填The function definition. Contains name (required), description, and parameters (a JSON Schema object), plus optional strict.
tool_choicestring | objectControls which (if any) tool is called by the model. none means no tool call (default when no tools are present), auto lets the model decide (default when tools are present), required forces one or more tool calls; you can also force a specific function with { "type": "function", "function": { "name": "my_function" } }.
parallel_tool_callsbooleanWhether to enable parallel function calling during tool use, officially defaults to true.
reasoning_effortstringConstrains effort on reasoning for reasoning models: none, minimal, low, medium, high, xhigh, max; officially defaults to medium. Lower effort yields faster responses and fewer reasoning tokens. Not all reasoning models support every value; effectiveness depends on the model.
verbositystringConstrains the verbosity of the model's response: low, medium, high; officially defaults to medium. Lower values produce more concise responses. Supported by GPT-5 family models.
modalitiesarray<string>Output types you would like the model to generate. Defaults to ["text"]; audio models (e.g. gpt-4o-audio-preview) can use ["text", "audio"].
audioobjectParameters for audio output. Required when modalities includes audio. Contains voice (built-in voices such as alloy, ash, coral, echo, nova, onyx, sage, shimmer, etc.) and format (wav, mp3, flac, opus, pcm16, aac).
predictionobjectConfiguration for a Predicted Output, which can greatly improve response times when large parts of the response are known ahead of time (e.g. regenerating a file with minor changes). Shape: { "type": "content", "content": ... }.
storebooleanWhether to store the output of this request (used upstream for distillation / evals products), officially defaults to false.
metadataobjectSet of up to 16 key-value pairs for storing additional structured information. Keys are strings up to 64 characters; values up to 512 characters.
service_tierstringThe processing tier used for serving the request: auto (default, uses the project setting), default, flex, scale, priority, fast. The service_tier in the response reflects the tier actually used and may differ from the request value. Effectiveness depends on the upstream.
web_search_optionsobjectWeb search options (for search-capable models). Supports search_context_size (low / medium / high) and user_location (approximate location).
prompt_cache_keystringA cache key used to optimize cache hit rates for similar requests; officially replaces the user field.
safety_identifierstringA stable identifier used to help detect end users who may be violating usage policies, up to 64 characters. Hashing the username or email is recommended to avoid sending identifying information.
userstringA stable identifier for your end users. Deprecated upstream in favor of safety_identifier and prompt_cache_key.
Request Examples
Response Examples
Response Fields
idstringThe response ID generated for this request.
objectstringThe streaming response is always chat.completion.chunk.
choicesarray<object>Array of candidate results.
deltaobjectIncremental content. May include role, content, reasoning_content, or tool_calls.
finish_reasonstringReason for completion. Common values are stop, length, and tool_calls.
usageobjectUsage statistics. It is guaranteed to appear only when the upstream returns usage data and stream_options.include_usage is enabled.