Seedance-2

Create Video

Submit a Seedance 2.0 multimodal video generation task using `POST /v1/video/generations`.

POSThttps://zx1.deepwl.net/v1/video/generations
Request
{
  "model": "doubao-seedance-2-0-260128",
  "content": [
    {
      "type": "text",
      "text": "First-person view fruit tea ad: 0-2s picking apples by hand; 2-4s cut to pouring into a shaker cup and shaking; 4-6s close-up of pouring into a clear cup; 6-8s raise the cup toward the camera"
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/apple.jpg" },
      "role": "reference_image"
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/cup.jpg" },
      "role": "reference_image"
    },
    {
      "type": "video_url",
      "video_url": { "url": "https://example.com/pov_reference.mp4" },
      "role": "reference_video"
    },
    {
      "type": "audio_url",
      "audio_url": { "url": "https://example.com/bgm.mp3" },
      "role": "reference_audio"
    }
  ],
  "metadata": {
    "duration": 8,
    "resolution": "720p",
    "ratio": "16:9",
    "generate_audio": true,
    "watermark": false
  }
}
Response
{
  "task_id": "cgt-20260412163502-x8k2m"
}

Submit a Seedance 2.0 video generation task. It supports text-to-video, first-frame/first-and-last-frame, reference image/video/audio, video continuation, video editing, and multimodal composition modes. For assets in the media library, it is recommended to reference them in content using asset://{assetId} (see Upload Assets).

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

Model name:

  • doubao-seedance-2-0-260128: Standard version, optimized for the best visual quality and complex shot planning
  • doubao-seedance-2-0-fast-260128: Fast version, optimized for low latency and cost-sensitive scenarios
contentarray<object>必填

Multimodal input array; the order affects role assignment.

typestring必填

Content type: text, image_url, video_url, audio_url, draft_task.

textstring

Required when type=text; prompt text.

image_urlobject

Used when type=image_url; must include url.

urlstring必填

Public image URL or asset reference asset://{assetId}.

video_urlobject

Used when type=video_url; must include url.

urlstring必填

Public video URL or asset://{assetId}.

audio_urlobject

Used when type=audio_url; must include url.

urlstring必填

Public audio URL or asset://{assetId}.

draft_taskobject

Used when type=draft_task; must include id, and it must be the only element in content.

idstring必填

Draft task ID, used to continue generation from a draft.

rolestring

Media role:

  • first_frame: first frame (image)
  • last_frame: last frame (image)
  • reference_image: reference image
  • reference_video: reference/source video (continuation, editing)
  • reference_audio: reference audio (requires metadata.generate_audio=true)
metadataobject

Video generation parameters; all are optional.

durationinteger

Video duration in seconds. Valid range [4, 15] or -1 (automatically determined by the model), default 5.

resolutionstring

Resolution: 480p, 720p, 1080p, 4k, default 720p.

ratiostring

Aspect ratio: 16:9, 9:16, 1:1, 4:3, adaptive, default 16:9.

framesinteger

Total number of video frames. Mutually exclusive with duration; if frames is provided, it takes precedence over duration.

seedinteger

Random seed. The same seed plus the same input can produce similar results.

camera_fixedboolean

Whether to keep the camera fixed (suppress camera movement), default false.

watermarkboolean

Whether to add a watermark in the bottom-right corner of the video, default true.

generate_audioboolean

Whether to generate or synthesize audio. Must be true when using reference_audio, default false.

return_last_frameboolean

Whether to return the final frame image URL for subsequent continuation, default false.

draftboolean

Draft mode: faster generation with slightly lower quality, suitable for previews, default false.

service_tierstring

Service tier, default default.

execution_expires_afterinteger

Maximum task execution time in seconds, range [3600, 259200] (1 hour to 3 days), default 172800.

callback_urlstring

Callback URL when the task is completed.

Request Examples

Response Examples

After submission, use Query Task to poll task_id.

Response Fields

task_idstring

Task ID, used for Query Task.

content Mixing Rules

Violating the following rules may return 400:

  • reference_image cannot appear together with first_frame / last_frame
  • audio_url cannot be the only input in content; it must be paired with at least an image or video
  • draft_task must be the only element in the content array

Generation Mode Comparison

ModeRequest Example Labelcontent Key Points
Text to VideoText to Videotext + optional reference image/video/audio
First-Frame Image to VideoFirst-Frame Image to Videotext + first_frame
First-and-Last-Frame Image to VideoFirst-and-Last-Frame Image to Videotext + first_frame + last_frame
Reference Image to VideoReference Image to Videotext + reference_image
Video ContinuationVideo Continuationtext + reference_video
Video EditingVideo Editingtext + reference_video + reference_image
Multimodal CompositionMultimodal Compositiontext + multiple types of references
Reference AssetsReference AssetsEach URL uses asset://{assetId}