Gemini Image Generation API
Call Gemini image models via `POST /v1beta/models/{model}:generateContent`, supporting text-to-image and image-to-image.
https://zx1.deepwl.net/v1beta/models/{model}:generateContentcurl -X POST https://zx1.deepwl.net/v1beta/models/gemini-3-pro-image-preview:generateContent \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"contents": [{
"role": "user",
"parts": [
{ "text": "A futuristic AI workspace, cinematic lighting" }
]
}],
"generationConfig": {
"responseModalities": ["TEXT", "IMAGE"],
"imageConfig": { "aspectRatio": "16:9", "imageSize": "2K" }
}
}'{
"candidates": [{
"content": {
"role": "model",
"parts": [
{ "inlineData": { "mimeType": "image/png", "data": "BASE64_OR_URL" } }
]
},
"finishReason": "STOP"
}],
"usageMetadata": {
"promptTokenCount": 123,
"candidatesTokenCount": 456,
"totalTokenCount": 579
},
"modelVersion": "gemini-3-pro-image-preview"
}Call Gemini image models through the official generateContent protocol. Supported models: gemini-3-pro-image-preview, gemini-2.5-flash-image-preview, gemini-3.1-flash-image-preview.
Text-to-image and image-to-image share the same endpoint — for image-to-image, simply append reference images to parts[]. See the "Text-to-Image / Image-to-Image" tabs on the right for complete request examples.
Request Headers
Authorizationstring必填Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY.
Content-Typestring必填Request content type, must be application/json.
Request Body
contentsarray<object>必填Input content array, usually containing a single element.
rolestringMessage role, fixed to user.
partsarray<object>必填Content parts. Text-to-image uses only text; for image-to-image, append one inlineData per reference image after the text part. Multiple reference images are supported.
textstringText prompt. For image-to-image, describe the desired edits, what to preserve, and the output style.
inlineDataobjectReference image (image-to-image only).
mimeTypestringMIME type of the reference image: image/jpeg, image/png, image/webp.
datastringBase64-encoded content of the reference image.
generationConfigobject必填Generation configuration.
responseModalitiesarray<string>必填Output modalities. Must include IMAGE for image generation; ["TEXT", "IMAGE"] is recommended.
imageConfigobjectImage configuration.
aspectRatiostringOutput aspect ratio: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9. Defaults to 1:1. gemini-3.1-flash-image-preview additionally supports the extreme ratios 1:4, 4:1, 1:8, 8:1.
imageSizestringOutput resolution: 1K, 2K, 4K. Defaults to 1K; some models also accept 512 (0.5K, no K suffix). Note: Flash Image channels such as gemini-3.1-flash-image-preview ignore this field and always return roughly 1K.
temperaturenumberSampling temperature, range 0 to 2. Defaults to 1.0.
topPnumberNucleus sampling parameter. Defaults to 0.95.
topKintegerTop-K sampling parameter.
maxOutputTokensintegerMaximum output tokens. Use at least 8192 for image generation.
seedintegerRandom seed. Reusing the same seed improves reproducibility.
safetySettingsarray<object>Optional safety filter settings. Each element: { "category": "HARM_CATEGORY_*", "threshold": "BLOCK_*" }.
Request Examples
Response Examples
The generated image is in inlineData under candidates[0].content.parts[]; data is either Base64-encoded content or an image URL.
Response Fields
candidatesarray<object>Array of candidate results.
finishReasonstringFinish reason. STOP means normal completion; SAFETY means blocked by safety policies.
content.parts[].inlineDataobjectGenerated image content.
mimeTypestringMIME type of the image, usually image/png.
datastringImage data — Base64-encoded content or an image URL. Clients must handle both forms.
usageMetadataobjectToken usage statistics.
promptTokenCountintegerNumber of input tokens.
candidatesTokenCountintegerNumber of output tokens.
totalTokenCountintegerTotal token count.