Skip to main content
Generate text responses with Google’s Gemini models. Provide a text prompt and, optionally, one or more images, audio clips, videos, or files as multimodal context.

Inputs

Common Inputs

Gemini 3.7 Flash Inputs

These inputs appear when model is set to "Gemini 3.7 Flash".

Gemini 3.5 Flash Inputs

These inputs appear when model is set to "Gemini 3.5 Flash".

Gemini 3.1 Pro Inputs

These inputs appear when model is set to "Gemini 3.1 Pro".

Gemini 3.1 Flash-Lite Inputs

These inputs appear when model is set to "Gemini 3.1 Flash-Lite".

Media and File Inputs

The following inputs are shared by all four models and appear alongside the model-specific inputs. Note: When media (images, audio, or video) is attached, the node uploads the first 10 media items to ComfyAPI storage and passes them as URLs; this URL budget is shared across all media types and is consumed in order (video first, then audio, then images). Any remaining media is encoded inline as base64 data, with a maximum combined inline payload of 18 MB. If the inline payload would exceed 18 MB, the node raises an error. The prompt parameter must contain at least one non-whitespace character. Setting seed to 0 requests a random seed.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 00e0f614303fa723eb787ad763e0b0c6322f89abf43d93b697357527b2fae49c