OpenAI Format

Realtime API

Use the OpenAI Realtime-compatible WebSocket route for real-time interactions.

GEThttps://zx1.deepwl.net/v1/realtime

The Realtime route uses a WebSocket connection and is suitable for low-latency voice conversations, real-time event streams, and bidirectional interaction scenarios. This route goes through TokenAuth, model rate limiting, and channel distribution.

Connection URL

wss://https://zx1.deepwl.net/v1/realtime

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY. If the client cannot conveniently set WebSocket request headers, you can use a supported temporary token or have the proxy layer inject the authentication header according to your gateway configuration.

Connection Example

const ws = new WebSocket("wss://https://zx1.deepwl.net/v1/realtime", [], {
  headers: {
    Authorization: "Bearer YOUR_API_KEY",
  },
});

ws.onopen = () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      modalities: ["text"],
      instructions: "You are a real-time voice assistant."
    }
  }));
};

ws.onmessage = (event) => {
  console.log(JSON.parse(event.data));
};

Session Configuration (session.update)

After the connection is established, the client can send a session.update event at any time to update the session configuration. The server responds with a session.updated event containing the full effective configuration. Only the fields present in the message are updated. To clear instructions, pass an empty string; to clear tools, pass an empty array; to disable turn_detection, pass null.

modalitiesarray<string>

The set of modalities the model can respond with. Defaults to ["text", "audio"]. Set to ["text"] to disable audio output.

modelstring

The Realtime model used for this session, e.g. gpt-realtime, gpt-4o-realtime-preview. It can also be specified via the model query parameter in the connection URL.

instructionsstring

The default system instructions for the session, used to guide the model's response content and format (e.g. "be extremely succinct") as well as audio behavior (e.g. "talk quickly").

voicestring默认值: alloy

The voice the model uses to respond. Built-in voices: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar. The voice cannot be changed once the model has responded with audio.

input_audio_formatstring默认值: pcm16

The format of the input audio: pcm16 (24kHz PCM), g711_ulaw, or g711_alaw.

output_audio_formatstring默认值: pcm16

The format of the output audio: pcm16, g711_ulaw, or g711_alaw.

input_audio_transcriptionobject

Configuration for input audio transcription, off by default.

modelstring

The transcription model, e.g. whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe.

languagestring

The language of the input audio (ISO-639-1 format, e.g. zh, en). Supplying it improves accuracy and latency.

promptstring

An optional prompt to guide the transcription style or provide domain-specific terms.

turn_detectionobject

Server-side turn detection configuration. Pass null to disable; when disabled, the client must trigger model responses manually.

typestring默认值: server_vad

Detection type: server_vad (volume-based voice activity detection) or semantic_vad (uses a model to semantically estimate whether the user has finished speaking; more natural, but may have higher latency).

thresholdnumber默认值: 0.5

server_vad only. Activation threshold for VAD (0.0 to 1.0). A higher threshold requires louder audio to activate and may perform better in noisy environments.

prefix_padding_msinteger默认值: 300

server_vad only. Amount of audio (in milliseconds) to include before the VAD-detected speech start.

silence_duration_msinteger默认值: 500

server_vad only. Duration of silence (in milliseconds) before speech is considered stopped. Shorter values make the model respond more quickly, but it may jump in on short pauses.

eagernessstring默认值: auto

semantic_vad only. How eagerly the model responds: low (waits longer for the user to continue), medium, high (responds more quickly), auto (default, equivalent to medium). low / medium / high have max timeouts of about 8s, 4s, and 2s respectively.

create_responseboolean默认值: true

Whether to automatically generate a response when a VAD stop event occurs.

interrupt_responseboolean默认值: true

Whether to automatically interrupt (cancel) any ongoing response when a VAD start event occurs. If both create_response and interrupt_response are false, the model never responds automatically, but VAD events are still emitted.

input_audio_noise_reductionobject

Configuration for input audio noise reduction, applied before the audio is sent to VAD and the model; filtering can improve turn detection accuracy. Pass null to disable.

typestring

Noise reduction type: near_field (close-talking microphones such as headsets) or far_field (far-field microphones such as laptop or conference room microphones).

temperaturenumber默认值: 1

Sampling temperature, between 0.6 and 1.2. Defaults to 1.

max_response_output_tokensinteger | string默认值: inf

Maximum number of output tokens for a single assistant response, inclusive of tool calls. Provide an integer between 1 and 4096 to limit output tokens, or "inf" for the maximum available tokens for the model (default).

toolsarray<object>

Tools available to the model. Function tools are currently supported: { "type": "function", "name": ..., "description": ..., "parameters": ...(JSON Schema) }.

tool_choicestring | object默认值: auto

How the model chooses tools: auto (default), none, required, or { "type": "function", "name": "..." } to force a specific function call.

Common Events

Client events:

eventDescription
session.updateUpdate session configuration; see fields above
input_audio_buffer.appendAppend Base64-encoded input audio
input_audio_buffer.commitCommit the input audio buffer (manual commit when VAD is off)
input_audio_buffer.clearClear the input audio buffer
conversation.item.createCreate a conversation item
conversation.item.deleteDelete a conversation item
response.createTrigger a model response
response.cancelCancel an in-progress response

Server events:

eventDescription
session.created / session.updatedSession created / configuration updated
input_audio_buffer.speech_started / speech_stoppedVAD detected speech start / stop
conversation.item.input_audio_transcription.completedInput audio transcription completed
response.created / response.doneResponse started / completed
response.text.deltaText delta
response.audio.delta / response.audio.doneAudio delta / audio output complete
response.audio_transcript.delta / doneOutput audio transcript delta / complete
response.function_call_arguments.doneFunction call arguments complete
errorError event
rate_limits.updatedRate limits updated