Qwen Image Image Editing API
Use `POST /v1/images/edits` to call `qwen-image-3.0` and `qwen-image-3.0-pro` for image-to-image/editing, supporting JSON multimodal requests and multipart file uploads.
https://zx1.deepwl.net/v1/images/editscurl --request POST 'https://zx1.deepwl.net/v1/images/edits' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"model": "qwen-image-3.0-pro",
"input": {
"messages": [
{
"role": "user",
"content": [
{
"image": "https://example.com/reference.png"
},
{
"text": "Keep the facial features, change the outfit to a modern professional style"
}
]
}
]
},
"parameters": {
"prompt_extend": true,
"size": "1080*1080",
"n": 1,
"watermark": false
}
}'{
"created": 1788188983,
"data": [
{
"url": "https://example.com/generated.png"
}
],
"metadata": {
"output": {
"choices": [
{
"finish_reason": "stop",
"message": {
"role": "assistant",
"content": [
{
"type": "image",
"image": "https://example.com/generated.png"
}
]
}
}
]
},
"request_id": "example-request-id",
"usage": {
"input_image_count": 1,
"input_image_type": "qima_input_1k",
"output_image_count": 1,
"output_image_type": "qima_output_1k",
"output_width": 1080,
"output_height": 1080
}
}
}Image editing shares the model family with text-to-image, using the editing route. Two submission methods are supported: JSON multimodal requests and multipart/form-data file uploads.
Request Headers
Authorizationstring必填Request authentication. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
Content-TypestringRequest content type. Must be application/json for JSON request bodies; for multipart/form-data it is carried automatically by the client (e.g. when curl uses -F), no need to set it manually.
Request Body
modelstring必填Must be qwen-image-3.0 or qwen-image-3.0-pro.
inputobjectThe native multimodal input structure containing messages.
messagesarray必填Message list; pass only one message.
rolestring必填Message role, fixed to user.
contentarray必填Message content list. Each content element is either text or image. Image-to-image/editing passes 1–3 image items and one text.
textstringText content, i.e. the edit instruction.
imagestringReference image address, supporting HTTP/HTTPS URLs or Data URLs with a MIME prefix. Supported formats: JPG, JPEG, PNG, BMP, TIFF, WEBP, GIF; officially recommended width and height between 384 and 2048 pixels; a single image must not exceed 10 MB.
parametersobjectQwen image generation parameters. When explicitly provided, the adapter uses these parameters directly and no longer copies size, n, or watermark from the top level.
sizestringOutput size in the native format WIDTH*HEIGHT (e.g. 1080*1080).
nintegerNumber of output images; the current upstream range is 1–6. When omitted or 0, it is treated as 1.
watermarkbooleanWhether to add a watermark.
prompt_extendbooleanWhether to enable prompt extension.
seedintegerRandom seed.
promptstringEdit instruction used in the multipart format.
imagefile | stringReference image file in the multipart format; can be repeated multiple times.
nintegerNumber of images in the multipart format. When omitted or 0, it is treated as 1.
watermarkbooleanWatermark switch in the multipart format.
multipart file uploads
Multipart requests submit the following fields with --form (see the Request Example panel on the right for a runnable example):
modelprompt- one or more
image nwatermark
Multiple files can be passed by repeating the image field:
--form 'image=@/absolute/path/reference-1.png' \
--form 'image=@/absolute/path/reference-2.png'
The multipart path does not build the full Qwen image parameters. When size, prompt_extend, or seed needs to be controlled, use the JSON multimodal request.
Request Example
Response Example
Response Fields
Image-to-image/editing shares the same response structure as text-to-image. The field definitions for created, data (with url, revised_prompt), metadata (with output, request_id, usage), and the error object (message, type, param, code), along with error-code meanings and troubleshooting, are documented in Image Generation API · Response Fields.
The only difference from text-to-image: metadata.usage.input_image_count reflects the number of reference images actually passed (0 for text-to-image).