AI NewsWords 1509Read time4 min

Google Launches Gemini Omni 1.1 Flash With 40-Second Scene Extension and 4K Upscaling

Gemini Omni 1.1 Flash adds longer scene extension, keyframe control, faster 360p drafts, video references, and 4K upscaling.

Google has released Gemini Omni 1.1 Flash, a production-ready update to its multimodal video-generation and editing model. The central upgrade is scene extension: the model can examine as much as 10 seconds of preceding footage when generating a continuation, compared with one second for Google’s previous models.

Developers can extend a video in 10-second increments to a cumulative length of 40 seconds. Google says the larger reference window helps the model preserve characters, motion, lighting, audio, and narrative context across successive extensions, addressing one of the practical barriers to producing sequences longer than a single generated shot.

The stable API model is identified as gemini-omni-1.1-flash. It accepts text, images, and video and produces video with audio. Google’s model listing specifies output durations from three to 10 seconds, resolutions of 360p, 720p, 1080p, or 4K, and a frame rate of 24 frames per second.

1. Scene Extension Now Uses Ten Seconds of Context

Scene extension generates new footage at the end of an existing clip. Gemini Omni 1.1 Flash can extend either a video previously generated in an API interaction or an uploaded video supplied through Google’s Files API.

For model-generated video, developers can pass the earlier interaction’s identifier through previous_interaction_id. This preserves the state of the prior generation without requiring the resulting video to be uploaded again. Developers can then request a continuation in natural language, including instructions for camera movement, dialogue, music, or events in the next part of the scene.

Uploaded clips can also be extended with a prompt, but the API documentation sets tighter conditions. An uploaded input must be no longer than 10 seconds, and the model can only append footage at the end. It cannot prepend a new opening or insert an extension into the middle of a clip.

The dialogue rules also depend on the workflow. Gemini Omni 1.1 Flash can generate additional spoken dialogue when extending its own earlier output through a multi-turn interaction. It cannot currently extend an uploaded video in which someone is already speaking and add more dialogue, although a silent continuation remains possible.

Editing or extending uploaded videos is unavailable through the API in the European Economic Area, Switzerland, and the United Kingdom. Extensions of videos generated by the model remain supported in available regions.

These restrictions matter because Google’s 40-second figure describes cumulative, multi-step extension rather than a single 40-second generation. Each continuation is still produced as a shorter segment, with the previous footage serving as context for the next one.

2. Keyframes and References Add More Directorial Control

Gemini Omni 1.1 Flash adds first-and-last-frame interpolation. A developer can supply two images representing the beginning and end of a shot, then describe the desired transition. The model generates the intervening motion as one continuous video.

That mechanism gives creators more control than a text prompt alone. The endpoints can constrain composition and subject placement while the prompt specifies how the camera or scene should move between them. Google presents camera orbits, zoom transitions, whip pans, and looping clips as intended uses.

The model also accepts short video references. Google says a reference can contribute as much as three seconds of motion, visual context, or character information to a new scene. The API documentation notes that audio in a video reference is ignored, so the feature should be understood as a visual and motion reference rather than an audio-transfer system.

These controls extend the original Gemini Omni approach rather than replace it. The first Omni Flash release already supported conversational editing: each natural-language revision could build on the previous output while attempting to preserve parts of the scene that the user did not ask to change. Version 1.1 adds more explicit constraints for shot boundaries, continuations, and reference-driven motion.

Stateful editing still depends on storing the interaction. If a developer disables storage, the resulting video cannot later be edited through its previous_interaction_id. Large outputs also require different handling: Google recommends URI delivery for video files over 4 MB rather than returning the entire file inline.

3. A Draft-to-Delivery Resolution Workflow

Google is positioning 360p generation as a drafting mode. The company says 360p previews can be generated up to 60% faster than standard 720p output, based on comparative system throughput, and at one-third of the cost.

The performance statement is specifically a comparison between 360p and 720p generation. It is not a general claim that every Omni 1.1 request will complete 60% faster than the previous model.

This lower-resolution option is designed for workflows in which several prompt, camera, or composition variants are tested before a final clip is selected. Applications can generate inexpensive previews, compare them, and reserve higher-resolution processing for the chosen result.

The default API resolution is 720p. Developers can request 1080p or 4K output, but Google’s documentation identifies both formats as upscaled output. The model is therefore not documented as generating native 4K frames directly; it produces an upscaled delivery version.

That distinction is relevant for production teams evaluating fine detail. A 4K file provides a higher-resolution deliverable, but the label alone does not establish the same detail retention as footage originally generated or captured at native 4K resolution.

4. Availability and the Shift From Preview to Stable API

Google released gemini-omni-1.1-flash as the stable, generally available version of the model on August 27, 2026. The earlier gemini-omni-flash-preview identifier remains listed as the preview version, while Google’s deprecation page says no shutdown date has been announced for the stable model.

Developers can access Omni 1.1 through the Gemini API in Google AI Studio. Enterprise customers can build with it through the Gemini Enterprise Agent Platform. This release status is consequential for companies that were reluctant to integrate a preview identifier into maintained products: the stable model name now provides a defined production target, although Google has not published a retirement date.

Gemini Omni 1.1 Flash is also available globally in Google Flow for Google AI Plus, Pro, and Ultra subscribers. Scene extension is available to subscribers on those three tiers in the Gemini app. Google’s announcement does not say that every new API control is exposed identically in each consumer interface.

Existing integrations broaden the model’s reach beyond Google’s own tools. Google says Gemini Omni Flash has been integrated into Adobe Firefly, while Figma Weave and Runway are among the creative platforms using the model. Those integrations are company-reported examples rather than independent performance evaluations.

5. Practical Benefits and Documented Limits

For creative-tool developers, the release supplies several stages of a conventional video workflow through one model: low-resolution ideation, endpoint-controlled shot generation, conversational revision, scene continuation, and high-resolution delivery. The API can therefore support iterative interfaces in which users branch from earlier generations instead of restarting every shot from a new prompt.

The 10-second extension context may reduce visible discontinuities, but it does not guarantee them away. Google’s model card states that complete consistency across edits, scenes with complex motion, and perfectly accurate text remain challenging. Claims of improved continuity should consequently be treated as relative to the prior one-second context window, not as a promise of seamless results in every sequence.

The API also does not support voice editing or uploaded audio references. It restricts editing of images containing certain recognizable people, and additional limits apply to images containing minors in the EEA, Switzerland, and the United Kingdom.

For content produced or edited inside the Gemini app, Google Flow, or YouTube, Google says it applies an imperceptible SynthID watermark and C2PA Content Credentials. Google’s published statement names those products specifically and should not be read as confirmation that every video returned directly through the Gemini API receives the same provenance treatment.

Frequently Asked Questions

What is the API model name?

The stable model identifier is gemini-omni-1.1-flash. Google also lists gemini-omni-flash-preview as the preview version.

Can Gemini Omni 1.1 Flash generate a 40-second video in one request?

Google describes extensions in 10-second increments up to a cumulative length of 40 seconds. Standard output clips remain between three and 10 seconds.

Does the model generate native 4K video?

Google’s API documentation describes 1080p and 4K as upscaled outputs. The default generation resolution is 720p.

Can it extend any uploaded video?

No. Uploaded videos must be 10 seconds or shorter, extension is limited to the end of the clip, and regional and dialogue restrictions apply.

Where is Gemini Omni 1.1 Flash available?

It is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google AI Plus, Pro, and Ultra subscribers can also access it through Google Flow, with scene extension offered in the Gemini app.

Sources

Share

Share this article