Unified Video Generation API

Submit asynchronous generation tasks using the platform's unified video task format, compatible with multiple video models and providers.

POSThttps://zx1.deepwl.net/v1/video/generations
Request
curl -X POST https://zx1.deepwl.net/v1/video/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '
{
  "model": "kling-v1",
  "prompt": "A golden retriever running on the beach, with the camera smoothly tracking it",
  "duration": 5,
  "size": "1280x720",
  "metadata": {
    "style": "cinematic"
  }
}'
Response
{
  "id": "video_task_abc123",
  "task_id": "video_task_abc123",
  "object": "video",
  "model": "kling-v1",
  "status": "queued",
  "progress": 0,
  "created_at": 1735689600
}

This is the entry point for multi-vendor video tasks. Requests first enter the unified task protocol, and then the dispatch layer selects the actual channel for execution.

  • A single entry point supports both text-to-video and image-to-video.
  • Asynchronous processing mode; after a successful submission, a public video task ID is returned.
  • Supports passing through vendor-specific parameters via metadata.
  • Can be used with /v1/video/generations/{task_id} and /v1/videos/{task_id} to form a complete task flow.

Request Headers

Authorizationstring必填

Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.

Content-Typestring必填

Request content type, must be application/json.

Request Body

modelstring必填

The model name used for billing and dispatch, such as kling-v1, sora-2, or veo-3. Whether it is ultimately available depends on the current Token's assigned group and channel configuration.

promptstring

Text prompt. It is usually required for pure text-to-video generation; if you are only making light edits based on an image, it is still recommended to provide a clear description.

imagestring

Input image URL or Base64 for image-to-video generation. Together with images and input_reference, this indicates that visual reference input is present.

imagesarray<string>

Multi-image reference input. Commonly used for upstreams that require multiple reference images.

input_referencestring | array<string>

Additional reference materials. Some upstreams treat this as the first frame, reference image, or a collection of materials.

durationinteger

Target duration in seconds. Some channels also accept seconds; if both are provided, the exact precedence is determined by the adapter.

secondsstring

Duration field in OpenAI video-compatible style. Common values include 5, 10, and 15.

sizestring

Size or ratio hint, such as 1280x720, 720x1280, or 16:9. Whether it is recognized depends on the specific channel.

metadataobject

Vendor-specific extension parameters, such as style, negative_prompt, camera_control, aspectRatio, and quality.

Request Examples

Response Examples

Response Fields

idstring

Public task ID on the platform. It can be used directly for subsequent queries and proxy downloads.

task_idstring

Task ID field retained for compatibility with legacy APIs, usually the same as id.

objectstring

Fixed to video.

statusstring

Initial status after submission; common values are queued or in_progress.

progressinteger

Task progress percentage. It is usually 0 when first submitted.

Notes

This route is a unified task entry point and does not guarantee that all channels accept the same set of parameters. Fields not recognized by the adapter may be ignored, or may only take effect for certain models.