OpenAI Format

General Chat Completions API (Default Non-Streaming)

Use the OpenAI Chat Completions-compatible format to initiate a conversation and return the full result in one response.

POSThttps://zx1.deepwl.net/v1/chat/completions
Request
curl -X POST https://zx1.deepwl.net/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "gpt-4o",
  "messages": [
    { "role": "system", "content": "You are an API documentation assistant." },
    { "role": "user", "content": "Generate a summary of the API description." }
  ],
  "stream": false
}'
Response
{
  "id": "chatcmpl_abc123",
  "object": "chat.completion",
  "created": 1735689600,
  "model": "gpt-4o",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "This API accepts unified conversation messages and generates a complete response in a single shot based on the model."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 24,
    "total_tokens": 55
  }
}

Suitable for background tasks, structured output, short Q&A, and scenarios where real-time display of the generation process is not required. When stream is omitted or set to false, the API returns a complete chat.completion object in a single response.

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

Model ID, e.g. gpt-4o. You can check the models available to the current API Key via the model list.

messagesarray<object>必填

A list of messages comprising the conversation so far, in chronological order. Different models support different message modalities (text, images, audio, etc.). Common roles are developer (formerly system), system, user, assistant, and tool.

rolestring必填

The role of the message author: developer, system, user, assistant, or tool.

contentstring | array必填

Message content. A string indicates plain text; an array indicates multimodal content parts, supporting text, image_url, input_audio, and file. For tool messages, content is the tool result; for assistant messages, content may be null when only tool calls are made.

namestring

An optional name for the participant, helping the model distinguish between speakers with the same role.

tool_callsarray<object>

Only for assistant messages: the tool calls generated by the model. Each item contains id, type (always function), and function (with name and arguments).

tool_call_idstring

Only for tool messages (required): the ID of the tool call this message is responding to.

streamboolean

Whether to stream the response, officially defaults to false. In this project's route, when omitted or set to false, the request is handled as non-streaming and a complete chat.completion object is returned in one response.

response_formatobject

An object specifying the format that the model must output. Commonly used for JSON output or Structured Outputs.

typestring

text (default), json_object (the older JSON mode, ensuring valid JSON output), or json_schema (Structured Outputs, ensuring the output matches the supplied JSON schema; preferred for models that support it).

json_schemaobject

Provided when type is json_schema. Contains name and schema (a JSON Schema object), plus optional description and strict.

toolsarray<object>

A list of tools the model may call.

typestring必填

The type of the tool. Function tools are always function; custom tools are also supported.

functionobject必填

The function definition. Contains name (required), description, and parameters (a JSON Schema object), plus optional strict.

tool_choicestring | object

Controls which (if any) tool is called by the model. none means no tool call (default when no tools are present), auto lets the model decide (default when tools are present), required forces one or more tool calls; you can also force a specific function with { "type": "function", "function": { "name": "my_function" } }.

parallel_tool_callsboolean

Whether to enable parallel function calling during tool use, officially defaults to true.

temperaturenumber

Sampling temperature, between 0 and 2, officially defaults to 1. Higher values make output more random, lower values more focused and deterministic. We generally recommend altering this or top_p but not both.

top_pnumber

Nucleus sampling parameter, between 0 and 1, officially defaults to 1. It is generally not recommended to adjust temperature and top_p significantly at the same time.

ninteger

How many chat completion choices to generate for each input message, between 1 and 128, officially defaults to 1. You are charged for tokens across all choices.

stopstring | array<string>

Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence. Not supported with the latest reasoning models such as o3 and o4-mini.

frequency_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the likelihood of repeating the same line verbatim.

presence_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the likelihood of talking about new topics.

logit_biasobject

Modifies the likelihood of specified tokens appearing in the completion. A JSON object mapping token IDs (in the tokenizer) to bias values from -100 to 100.

logprobsboolean

Whether to return log probabilities of the output tokens, officially defaults to false.

top_logprobsinteger

An integer between 0 and 20 specifying the number of most likely tokens to return at each token position, each with its log probability. logprobs must be set to true when using this parameter.

max_completion_tokensinteger

An upper bound for the number of tokens that can be generated, including visible output tokens and reasoning tokens. Recommended for reasoning models.

max_tokensinteger

The maximum number of tokens that can be generated. Deprecated upstream in favor of max_completion_tokens, and not compatible with o-series reasoning models.

seedinteger

Random seed (Beta). With the same seed and parameters, the system makes a best effort to sample deterministically, though determinism is not guaranteed; when supported by the upstream this improves reproducibility. Monitor backend changes via the system_fingerprint response field.

reasoning_effortstring

Constrains effort on reasoning for reasoning models: none, minimal, low, medium, high, xhigh, max; officially defaults to medium. Not all reasoning models support every value; effectiveness depends on the model.

verbositystring

Constrains the verbosity of the model's response: low, medium, high; officially defaults to medium. Supported by GPT-5 family models.

modalitiesarray<string>

Output types you would like the model to generate. Defaults to ["text"]; audio models can use ["text", "audio"].

audioobject

Parameters for audio output. Required when modalities includes audio. Contains voice (e.g. alloy, ash, coral, echo, nova, onyx, sage, shimmer) and format (wav, mp3, flac, opus, pcm16, aac).

predictionobject

Configuration for a Predicted Output, which can greatly improve response times when large parts of the response are known ahead of time. Shape: { "type": "content", "content": ... }.

storeboolean

Whether to store the output of this request (used upstream for distillation / evals products), officially defaults to false.

metadataobject

Set of up to 16 key-value pairs for storing additional structured information. Keys are strings up to 64 characters; values up to 512 characters.

service_tierstring

The processing tier used for serving the request: auto (default), default, flex, scale, priority, fast. Effectiveness depends on the upstream.

web_search_optionsobject

Web search options (for search-capable models). Supports search_context_size (low / medium / high) and user_location.

prompt_cache_keystring

A cache key used to optimize cache hit rates for similar requests; officially replaces the user field.

safety_identifierstring

A stable identifier used to help detect end users who may be violating usage policies, up to 64 characters. Hashing the username or email is recommended.

userstring

A stable identifier for your end users. Deprecated upstream in favor of safety_identifier and prompt_cache_key.

Request Examples

Response Examples

Response Fields

choicesarray<object>

Array of candidate results.

messageobject

The message generated by the model.

contentstring | null

The text content generated by the model. May be null when a tool call occurs.

tool_callsarray<object>

The function tools requested by the model.

usageobject

Token usage for this request.