OpenAI Format

Text Completion API

Generate text using the OpenAI Legacy Completions compatible format.

POSThttps://zx1.deepwl.net/v1/completions
Request
curl -X POST https://zx1.deepwl.net/v1/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "gpt-3.5-turbo-instruct",
  "prompt": "把下面一句话改写得更正式:这个接口挺好用。",
  "max_tokens": 120,
  "temperature": 0.3
}'
Response
{
  "id": "cmpl_abc123",
  "object": "text_completion",
  "created": 1735689600,
  "model": "gpt-3.5-turbo-instruct",
  "choices": [
    {
      "index": 0,
      "text": "该接口具备良好的易用性。",
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 22,
    "completion_tokens": 10,
    "total_tokens": 32
  }
}

The Legacy Completions API uses prompt as input and is suitable for businesses still using the old OpenAI SDK or simple text completion workflows. For new conversational scenarios, it is recommended to use the chat completions API.

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

Model ID, e.g. gpt-3.5-turbo-instruct. You can check available models via the Model List.

promptstring | array必填

The prompt(s) to generate completions for. Can be a string, an array of strings, an array of tokens, or an array of token arrays. Officially defaults to <|endoftext|> (the model generates as if from the beginning of a new document).

suffixstring

The suffix that comes after a completion of inserted text. Only supported for gpt-3.5-turbo-instruct.

max_tokensinteger

The maximum number of tokens to generate, officially defaults to 16. The token count of your prompt plus max_tokens cannot exceed the model's context length.

temperaturenumber

Sampling temperature, between 0 and 2, officially defaults to 1. Higher values make output more random, lower values more focused and deterministic. We generally recommend altering this or top_p but not both.

top_pnumber

Nucleus sampling parameter, between 0 and 1, officially defaults to 1.

ninteger

How many completions to generate for each prompt, between 1 and 128, officially defaults to 1. This can quickly consume your token quota; use reasonable max_tokens and stop settings.

streamboolean

Whether to stream back partial progress, officially defaults to false. When enabled, tokens are sent as server-sent events as they become available, terminated by data: [DONE].

stream_optionsobject

Additional options for streaming output. Only set this when stream: true.

include_usageboolean

Streams an extra chunk containing usage before data: [DONE], reporting token usage for the entire request.

logprobsinteger

Include the log probabilities on the logprobs most likely tokens, as well as the chosen tokens. Maximum value is 5; the response may contain up to logprobs + 1 elements.

echoboolean

Whether to echo back the prompt in addition to the completion, officially defaults to false.

stopstring | array<string>

Up to 4 sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence.

presence_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the likelihood of talking about new topics.

frequency_penaltynumber

Number between -2.0 and 2.0, officially defaults to 0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the likelihood of repeating the same line verbatim.

best_ofinteger

Generates best_of completions server-side and returns the one with the highest log probability per token. Officially defaults to 1, maximum 20. Cannot be streamed; when used with n, best_of must be greater than n. Use carefully as it can quickly consume your token quota.

logit_biasobject

Modifies the likelihood of specified tokens appearing in the completion. A JSON object mapping token IDs (in the GPT tokenizer) to bias values from -100 to 100. For example, passing {"50256": -100} prevents the <|endoftext|> token from being generated.

seedinteger

Random seed. With the same seed and parameters, the system makes a best effort to sample deterministically, though determinism is not guaranteed; monitor backend changes via the system_fingerprint response field.

userstring

A unique identifier representing your end-user, which can help monitor and detect abuse.

Request Examples

Response Examples

Response Fields

idstring

The completion response ID.

objectstring

Object type, always text_completion.

createdinteger

The timestamp when the response was created.

modelstring

The model that generated the response.

choicesarray<object>

Array of candidate results.

indexinteger

Index of the candidate result.

textstring

The text content generated by the model.

finish_reasonstring

Reason for completion, e.g. stop.

usageobject

Token usage statistics.

prompt_tokensinteger

Number of input tokens.

completion_tokensinteger

Number of output tokens.

total_tokensinteger

Total token count.