Qwen Image Image Generation API
Use `POST /v1/images/generations` to call `qwen-image-3.0` and `qwen-image-3.0-pro` for image generation, supporting both the native multimodal structure and the OpenAI simplified structure.
https://zx1.deepwl.net/v1/images/generationscurl --request POST 'https://zx1.deepwl.net/v1/images/generations' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"model": "qwen-image-3.0",
"input": {
"messages": [
{
"role": "user",
"content": [
{
"text": "A warm film-style street portrait in the afternoon, a person holding a bouquet of roses, backlit, shallow depth of field"
}
]
}
]
},
"parameters": {
"prompt_extend": true,
"size": "1080*1080",
"n": 1,
"watermark": false,
"seed": 123456
}
}'{
"created": 1788188983,
"data": [
{
"url": "https://example.com/generated.png"
}
],
"metadata": {
"output": {
"choices": [
{
"finish_reason": "stop",
"message": {
"role": "assistant",
"content": [
{
"type": "image",
"image": "https://example.com/generated.png"
}
]
}
}
]
},
"request_id": "example-request-id",
"usage": {
"input_image_count": 0,
"input_image_type": "qima_input_1k",
"output_image_count": 1,
"output_image_type": "qima_output_1k",
"output_width": 1080,
"output_height": 1080
}
}
}qwen-image-3.0 and qwen-image-3.0-pro use the unified text-to-image entry point (POST /v1/images/generations).
- When
prompt_extendorseedis needed, passing the native multimodal structure directly is recommended. - When only basic parameters are needed, the OpenAI simplified structure can also be used.
Request Headers
Authorizationstring必填Request authentication. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
Content-Typestring必填Request content type, must be application/json.
Request Body
modelstring必填Must be qwen-image-3.0 or qwen-image-3.0-pro.
promptstringGeneration prompt used in the simplified format; optional when input.messages is provided.
inputobjectThe native multimodal input structure containing messages. When provided, the top-level prompt can be empty; at least one valid text must be provided between input and the top-level prompt.
messagesarray必填Message list; pass only one message.
rolestring必填Message role, fixed to user.
contentarray必填Message content list. Each content element is either text or image. Text-to-image passes only one text; image-to-image/editing passes 1–3 image items and one text.
textstringText content, i.e. the generation prompt or edit instruction. Text-to-image must pass exactly one.
imagestringReference image address, supporting HTTP/HTTPS URLs or Data URLs with a MIME prefix. Image-to-image/editing passes 1–3; supported formats: JPG, JPEG, PNG, BMP, TIFF, WEBP, GIF; officially recommended width and height between 384 and 2048 pixels; a single image must not exceed 10 MB.
parametersobjectQwen image generation parameters. When explicitly provided, the adapter uses these parameters directly and no longer copies size, n, or watermark from the top level.
sizestringOutput size in the native format WIDTH*HEIGHT (e.g. 1080*1080). The simplified format uses the top-level size in WIDTHxHEIGHT.
nintegerNumber of output images; the current upstream range is 1–6. When omitted or 0, it is treated as 1.
watermarkbooleanWhether to add a watermark.
prompt_extendbooleanWhether to enable prompt extension.
seedintegerRandom seed.
nintegerNumber of images in the simplified format. When omitted or 0, it is treated as 1.
sizestringOutput size in the simplified format, using WIDTHxHEIGHT (e.g. 1080x1080).
watermarkbooleanWatermark switch in the simplified format.
OpenAI simplified structure
When only basic parameters are needed, prompt, size, n, and watermark can be passed at the top level:
{
"model": "qwen-image-3.0",
"prompt": "A warm film-style street portrait in the afternoon",
"size": "1080x1080",
"n": 1,
"watermark": false
}
The adapter performs the following conversions:
- Converts
prompttoinput.messages[0].content[0].text. - Converts the
xinsizeto the native format*. - When
nis omitted or0, setsnto1.
Request Example
Response Example
Response Fields
createdintegerUnix timestamp of when AIX processed the request.
dataarray<object>OpenAI-style array of image results.
urlstringImage URL returned by the upstream.
revised_promptstringAppears when the upstream returns text content.
metadataobjectThe original upstream response.
outputobjectThe original upstream output structure, where choices[].message.content[].image is the actual returned image address. When data is empty, combine this field with request_id to verify whether the upstream returned an image.
request_idstringUpstream request ID.
usageobjectInput/output image counts and resolution information.
errorobjectError structure returned when the request fails (e.g. HTTP 400).
messagestringError description.
typestringError type, currently fixed to ali_error.
paramstringThe failing parameter name, or an empty string if none.
codestringError code.