Kling Video

Kling Video Overview

Organize Kling's OpenAI-format interfaces by model series.

Kling Videos

The Kling video series uses the OpenAI-compatible /v1/videos interface and supports text-to-video, image-to-video, multiple reference images, first-and-last-frame, motion control, digital avatars, and lip sync.

OpenAI Format

PurposeMethodPath
Submit a video generation taskPOST/v1/videos
Query task status and resultsGET/v1/videos/{task_id}
Fetch the generated video contentGET/v1/videos/{task_id}/content

Supported Models

ModelDescriptionDuration Options (seconds)
Kling-3.0-OmniKling 3.0 all-in-one edition5, 10, 15 (default 5)
Kling-3.0Kling 3.05, 10 (default 5)
Kling-2.6Kling 2.65, 10 (default 5)
Kling-2.5Kling 2.55, 10 (default 5)
Kling-2.1Kling 2.15, 10 (default 5)
Kling-2.0Kling 2.05, 10 (default 5)
Kling-1.6Kling 1.65, 10 (default 5)
Kling-O1Kling O15, 10 (default 5)

Composite billing model names can also be passed directly as model; the gateway restores the upstream model name and fills in the parameters: kling-3.0-omni-1080p-ref-audio (1080P · with reference · with sound), kling-2.6-motion-pro-1080p (motion control · pro tier), kling-avatar-720p (digital avatar), kling-identify-face (lip sync). Kling-O3 and Kling-Mini are not preset yet.

Scene Types

Scene TypeDescriptionKey Conditions
motion_controlMotion controlA video reference is required
avatar_i2vDigital avatarBilled as the avatar scene (kling-avatar-720p)
lip_syncLip syncBilled at 5 seconds if shorter (kling-identify-face)
template_effectTemplate effectsVidu scene

Parameter Description

Basic Parameters (Top Level)

FieldTypeDescription
modelstringModel name (base model or composite billing model)
promptstringPrompt
secondsstring | integerVideo duration in seconds; top level takes priority; default 5
durationstring | integerDuration compatibility field; effective when top-level seconds is absent
sizestringQuick size (720P / 1080P or WxH)
imagestringSingle reference / first-frame image (URL or file ID)
imagesarray<string>Multiple reference images (up to 3)

metadata Parameters

FieldTypeDescription
output_configobjectOutput configuration (maps to upstream AigcVideoOutputConfig)
scene_typestringScene type (motion_control / avatar_i2v / lip_sync / template_effect)
motion_levelstringMotion control tier std / pro (billing tier)
offpeakbooleanWhether off-peak billing is enabled
last_frame_urlstringLast frame of a start/end frame pair (URL)
last_frame_file_idstringLast frame of a start/end frame pair (file ID)
video_urlstringReference video URL (required for motion control)
file_infosarray<object>Native FileInfos passthrough (up to 3 items)
ext_infostringNative ExtInfo string passthrough

output_config Parameters

FieldTypeDescription
resolutionstringResolution 720P / 1080P, default 720P
aspect_ratiostringAspect ratio 16:9 / 9:16 / 1:1 (default 16:9 for text-to-video)
durationfloatGeneration duration (seconds); Kling supports 5 / 10
audio_generationstringEnabled / Disabled
person_generationstringAllowAdult / Disallowed
input_compliance_checkstringEnabled / Disabled
output_compliance_checkstringEnabled / Disabled
enhance_switchstringEnabled / Disabled
storage_modestringPermanent / Temporary, default Temporary
media_namestringOutput media name (max 64 characters)
class_idintegerClass ID, default 0
expire_timestringExpiration time (ISO 8601)
frame_interpolatestringEnabled / Disabled (Vidu)
logo_addstringEnabled / Disabled (Vidu)

Parameter Rules

  • Duration priority: top-level seconds > top-level duration > metadata.seconds / metadata.duration / metadata.video_duration > default 5
  • Resolution priority: output_config.resolution > top-level size > default 720P
  • motion_control requires a video reference (video_url or a file_infos entry with Category=Video)
  • FileInfos is limited to 3 items; top-level images do not support base64 data URIs
  • lip_sync is billed at 5 seconds when shorter than 5 seconds
FeatureDescription
Motion controlscene_type: motion_control, precisely control motion from a reference video; tiers std / pro
Digital avatarscene_type: avatar_i2v, generate digital avatar videos
Lip syncscene_type: lip_sync, synchronized lip syncing with audio and video
Off-peak billingoffpeak: true, generate during off-peak hours for lower cost
Start/end framesimage + last_frame_url, control the start and end frames
Multiple reference imagesTop-level images (≤3) or native file_infos passthrough