Gemini Native Format
Use Google Gemini native paths and request bodies to call generateContent, streamGenerateContent, and model queries.
https://zx1.deepwl.net/v1beta/models/{model}:{action}curl -X POST https://zx1.deepwl.net/v1beta/models/gemini-2.0-flash:generateContent \
-H "x-goog-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '
{
"contents": [
{
"role": "user",
"parts": [
{ "text": "Introduce the Gemini native API in three sentences." }
]
}
],
"generationConfig": {
"temperature": 0.7,
"maxOutputTokens": 300
}
}'{
"candidates": [
{
"content": {
"role": "model",
"parts": [
{
"text": "The Gemini native API uses contents and parts to express input content. It supports text, images, files, and function calling. With a unified gateway, you can continue using the same API key and billing system."
}
]
},
"finishReason": "STOP",
"safetyRatings": [
{
"category": "HARM_CATEGORY_HARASSMENT",
"probability": "NEGLIGIBLE"
}
]
}
],
"usageMetadata": {
"promptTokenCount": 18,
"candidatesTokenCount": 58,
"totalTokenCount": 76
}
}Gemini Native Format preserves the Google Gemini API paths and request bodies. It is suitable for business integrations that already use the Gemini SDK, the contents/parts structure, or safety settings configuration.
Request Headers
Authorizationstring必填Authentication header. Uses a Bearer token, e.g. Bearer YOUR_API_KEY. Alternatively, you can use the x-goog-api-key header or the key query parameter.
x-goog-api-keystringGoogle API Key style authentication, e.g. YOUR_API_KEY. Use either this or Authorization.
Content-Typestring必填Request content type, must be application/json.
keystringGoogle API Key style query parameter authentication, e.g. /v1beta/models/gemini-2.0-flash:generateContent?key=YOUR_API_KEY.
Paths
| Method | Path | Description |
|---|---|---|
GET | /v1beta/models | Gemini model list |
POST | /v1beta/models/{model}:generateContent | Non-streaming content generation |
POST | /v1beta/models/{model}:streamGenerateContent | Streaming content generation |
Request Body
contentsarray<object>必填List of conversation contents. In multi-turn conversations, role alternates between user and model.
rolestringContent role: user or model.
partsarray<object>必填Content parts. Supports text, inlineData (mimeType + Base64 data), fileData (mimeType + fileUri), functionCall, functionResponse, and more.
videoMetadataobjectVideo input metadata. Use it on the same part as video inlineData or fileData to control the time range and frame sampling rate the model reads.
Supported fields:
startOffset: Video start offset, using Duration string format, such as"3s"or"3.5s".endOffset: Video end offset, using Duration string format, such as"10s".fps: Video sampling frame rate. Defaults to1.0; valid range is0 < fps <= 24.
When a request includes multiple videos, each video part can set its own videoMetadata.
systemInstructionobjectSystem-level instruction (System Prompt). Used to define model behavior, role setting, and response style. Compatible with both systemInstruction and system_instruction forms.
generationConfigobjectGeneration configuration, used to control model output behavior.
temperaturenumberControls output randomness, range 0 to 2, defaults to 1.0. Higher values produce more diverse results.
topPnumberNucleus Sampling probability threshold.
topKintegerSamples only from the top K most probable tokens.
candidateCountintegerNumber of candidate results to return. Defaults to 1.
maxOutputTokensintegerMaximum number of output tokens.
stopSequencesarray<string>Stops generation when the specified strings are encountered. Up to 5 sequences.
responseMimeTypestringSpecifies the output format: text/plain (default), application/json (JSON mode), text/x.enum.
responseSchemaobjectSchema constraint for structured output (a subset of OpenAPI 3.0 Schema). Must be used together with responseMimeType: application/json.
responseJsonSchemaobjectStructured output constraint in standard JSON Schema form. Mutually exclusive with responseSchema; also requires responseMimeType: application/json.
responseModalitiesarray<string>Output modalities, e.g. ["TEXT"]; image generation models use ["TEXT", "IMAGE"].
seedintegerFixes the random seed for reproducible results.
presencePenaltynumberReduces repetitive topics and encourages new content generation.
frequencyPenaltynumberReduces repetitive words or sentences.
responseLogprobsbooleanWhether to return token probability information.
logprobsintegerNumber of token probabilities to return, range 0 to 20.
enableEnhancedCivicAnswersbooleanWhether to enable enhanced answers for civic-related queries such as elections and government information.
speechConfigobjectSpeech output configuration (TTS), including voiceConfig, languageCode, and multi-speaker settings.
audioTimestampbooleanWhether to return audio timestamps. Only used for audio understanding scenarios.
thinkingConfigobjectReasoning configuration for Gemini Thinking models.
includeThoughtsbooleanWhether to include thought summaries in the response.
thinkingBudgetintegerThinking token budget (Gemini 2.5 series). -1 enables dynamic thinking; 0 disables thinking.
thinkingLevelstringThinking intensity level: low, high (used by the Gemini 3 series; mutually exclusive with thinkingBudget).
mediaResolutionstringMedia input resolution: MEDIA_RESOLUTION_LOW, MEDIA_RESOLUTION_MEDIUM, MEDIA_RESOLUTION_HIGH. Affects token consumption and understanding fidelity for images and videos.
imageConfigobjectImage generation configuration (image generation models only). Common subfields: aspectRatio (e.g. 1:1, 16:9), imageSize (1K, 2K, 4K).
safetySettingsarray<object>Safety policy configuration.
categorystring必填Risk category: HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_SEXUALLY_EXPLICIT, HARM_CATEGORY_DANGEROUS_CONTENT.
thresholdstring必填Risk blocking level: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE.
methodstringScoring method: SEVERITY (by severity) or PROBABILITY (by probability).
toolsarray<object>Gemini tool declaration list, used to enable function calling, search, code execution, and other capabilities. Supports the following tool types:
functionDeclarations: Function Calling function declarations.googleSearch: Google search capability.codeExecution: Code execution capability.urlContext: URL content parsing capability.retrieval: Retrieval-Augmented Generation (RAG). Deprecated; only available on older models.googleSearchRetrieval: Google retrieval augmentation. UsegoogleSearchfor newer models.
toolConfigobjectTool invocation configuration.
Commonly used to configure functionCallingConfig.mode: AUTO (model automatically decides whether to call tools), ANY (force tool calling), NONE (disable tool calling).
allowedFunctionNames: Specifies the list of functions allowed to be called.
cachedContentstringGemini Cached Content identifier. Used to reuse context cache, reducing token consumption and response latency for long-context requests.
labelsobjectCustom key-value labels for request tracking and billing attribution.
Request Examples
Put videoMetadata on the same part as the video data. It can be used to analyze only a video segment or adjust frame sampling density (see the "Video Input" tab).
Response Examples
Common Safety Settings
| category | Description |
|---|---|
HARM_CATEGORY_HARASSMENT | Harassment content |
HARM_CATEGORY_HATE_SPEECH | Hate speech |
HARM_CATEGORY_SEXUALLY_EXPLICIT | Sexually explicit content |
HARM_CATEGORY_DANGEROUS_CONTENT | Dangerous content |
| threshold | Description |
|---|---|
BLOCK_NONE | Do not block |
BLOCK_ONLY_HIGH | Block high risk only |
BLOCK_MEDIUM_AND_ABOVE | Block medium risk and above |
BLOCK_LOW_AND_ABOVE | Block low risk and above |