New: Gemini Omni 1.1 Flash is available

Gemini Omni 1.1 Flash AI Video Generator

Create multimodal AI videos with text, images, audio, characters, video references, or first and last frames in resolutions up to 4K.

One Fast Model, Two Clear Ways to Direct the Shot

Use multimodal references when the model should borrow context from source material, or use first and last frames when the transition itself matters. Keeping those paths separate prevents conflicting instructions.

Multimodal References

Combine images, audio IDs, character IDs, and one trimmed source video within the model's input budget.

First and Last Frames

Lock the opening frame and optionally pair it with an ending frame to shape a visual transition.

Flexible Output

Choose 4, 6, 8, or 10 seconds, landscape or portrait, and resolutions from 360p through 4K.

Repeatable Testing

Set a seed to make controlled prompt and reference comparisons easier, while allowing for model variation.

Build the Shot Around the Material You Already Have

Gemini Omni 1.1 Flash adapts to different production starting points without forcing every idea into the same prompt-only workflow.

Add up to seven images when no video or character references consume the shared seven-unit input budget.

Gemini Omni 1.1 Flash Capabilities

Practical controls for fast drafts, reference-led motion, and delivery-ready resolution tests.

Prompt to Video

Describe visual content, style, camera language, and character action in prompts up to 20,000 characters.

Image to Video

Use multiple public image references for subjects, scenes, styles, layouts, or storyboard continuity.

First-to-Last Frame

Define a starting image and optionally an ending image for a more intentional visual transition.

Video Reference

Guide the generation with one trimmed source clip while the model determines the resulting duration.

Audio and Characters

Bring in compatible audio and character assets when identity or sound direction needs more control.

360p to 4K

Test quickly at a lower resolution, then choose 720p, 1080p, or 4K when the direction is ready.

Gemini Omni 1.1 Flash FAQ

Key details about inputs, output controls, and model constraints.

What is Gemini Omni 1.1 Flash?

Gemini Omni 1.1 Flash is a multimodal AI video model available in the Gemini Omni workspace. It can generate video from a prompt plus supported image, audio, character, video, or frame guidance.


Which video durations are available?

Without a video input, you can choose 4, 6, 8, or 10 seconds. When a source video is used, the model determines output duration and ignores the duration setting.


Which aspect ratios and resolutions are supported?

The model supports 16:9 landscape and 9:16 portrait output at 360p, 720p, 1080p, or 4K.


Can I use the first frame with other references?

No. First-frame mode cannot be combined with image references, audio IDs, source videos, or character IDs. Use either the frame-guided path or the multimodal reference path for a generation.


Can I provide only a last frame?

No. A last frame must be paired with a first frame. The first frame can be used on its own.


How many references can I add?

Multimodal inputs share a seven-unit budget: each image uses one unit, a video uses two, and each character ID uses one. A request supports at most one video and up to three character IDs.


What are the source file limits?

Each image can be up to 20 MB. A source video can be up to 100 MB and 30 seconds long, with a selected segment no longer than 10 seconds.


How do I start using the model?

Open the video generator, select Gemini Omni 1.1 Flash, choose a compatible input path, set the output options, and submit the task.


Turn Your References Into a Focused Video Draft

Open the Gemini Omni video workspace, select 1.1 Flash, and build the next shot around the inputs that matter.