Create videos and edit scenes with text and reference assets
omni-flash is the video creation entry point for Google Gemini Omni Flash, suitable for creating new shots from text concepts, still images, or existing videos. It can generate visuals from prompts and use reference assets to adjust scenes, styles, and elements. Through this platform, you can select landscape or portrait aspect ratios and output resolutions, and integrate it into asset production workflows through asynchronous tasks.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation methods
Text-to-video, reference image guidance, video editing / video reference
Output aspect ratio
16:9, 9:16; default 16:9
Output resolution
720p, 1080p; default 720p
Image references
Submit one or more image links through image_urls
Video references
Up to 1 video link; at least 1 reference image must also be provided
Task delivery
Supports asynchronous queries and completion callbacks; returns video_url upon success
The specifications above are the omni-flash usage scope on this platform; extended capabilities of different native Gemini Omni versions should be distinguished separately.
Core capabilities
Turn shot intent into actionable descriptions
When starting with text, you can specify the subject, scene, action, camera movement, and lighting mood in the prompt to generate video visuals for review. Compared with writing only a theme, this is better for fully expressing filming intent: for example, the subject slowly turns around, the camera pushes forward, and the environment maintains soft backlighting, bringing creative discussions down to specific shots.
Bring static assets into video creation
With reference image-guided generation, you can continue creating dynamic visuals from product photos, character images, or illustrations. Images provide visual grounding, while text describes motion direction, scene relationships, and the desired style. This is suitable for production stages where design assets already exist but video shots have not yet been formed; clearly state whether the image serves as a subject reference or an overall visual reference.
Define modification goals around the original video
Existing videos can also serve as a starting point for creation, using reference images and editing instructions to generate new videos. You can request changes to the season, adjustments to the scene style, or the addition and removal of visual elements, while clearly specifying which compositional relationships should be preserved. This workflow is suitable for creating creative variations around the same asset rather than describing the entire scene from scratch each time.
Use Cases
Turn Product Images into Dynamic Presentations
Input selected product images, describe the product motion, camera direction, and background atmosphere, and create presentation videos for the marketing team to review. Landscape format is suitable for pages or presentation screens, while portrait format is suitable for vertical content layouts. After delivery, focus on checking whether the product outline, composition, and motion match the creative vision before deciding whether to proceed to final editing.
Visual Variations of the Same Scene
Input an existing scene video and reference images, use text to request changes to the weather, season, or visual style, and list the positional relationships that need to be preserved. For example, change a sunny beach into a winter scene while retaining the layout of the trees and boats. The output can be used to compare new versions across different creative directions, reducing the work of reimagining from a blank frame.
Integrate into Automated Asset Production
The application backend submits prompts, aspect ratios, and asset links, creates tasks with async, saves the task_id and then queries results, or receives completion notifications through callback_url. After obtaining a successful status and video link, download and archive it, then send it to review or editing workflows. This is suitable for production applications that do not want to keep generation request connections continuously open.
How to Choose This Model
Choose a Creation Method Based on Existing Assets
Choose text-to-video when you only have creative text; add image references when the subject appearance or visual direction has already been determined; use a combination of video and images when you want to modify existing shots. The value of choosing omni-flash is that it brings these asset formats into the same video workflow. For editing tasks especially, clearly state separately “what to change” and “what to preserve,” rather than submitting only vague style terms.
Differentiate the Video Endpoint from Native Versions
omni-flash is for video generation and editing and should not be confused with Gemini Flash text conversation models. Gemini Omni 1.1 Flash also has a separate native version name, so do not apply first-and-last-frame or extension controls directly to this endpoint just because the names are similar. If your goal is guidance from existing assets and scene rewriting, choose the workflow here; if you depend on specific version features, select models separately by version.
Get Started
Distinguish Generation from Adapting Existing Videos
Prepare a prompt for text generation; provide image_urls for image guidance. Video editing can pass at most one video_urls, and must also pass at least one reference image.
Set Up a Gemini Omni Video Request
Specify model=omni-flash and prompt to /gemini/videos, choose 16:9/9:16 and 720p/1080p, and describe what needs to be changed and preserved.
Save Completed Results
Save the task_id asynchronously, wait for succeeded through /gemini/tasks or a callback, then read video_url; download the finished video promptly and review the subject and scene changes segment by segment.
Trial Recommendation: Reference Image-Guided Scene Adaptation
Input and Goal
Keep the person and camera position in the original video consistent, replace the room wall style with the new background defined by the reference image, and use warm lighting overall.
Acceptance and Next Steps
Provide at most one video_urls and at least one image_urls, and describe the areas to modify and preserve; first check the subject and layout, then retrieve the completed video.
Usage Limits
Video editing is not as simple as submitting only a video: video_urls accepts at most one link, and at least one image_urls must be provided at the same time. The reference image should be relevant to the intended modification, and the prompt should clearly specify the editing effect; missing images will cause a request parameter error and prevent entry into the normal editing workflow.
Output resolution can be 720p or 1080p, but selecting a resolution does not guarantee that subject details and actions will meet requirements. Especially for scene replacement or adding and removing elements, review each item of the layout and appearance that needs to be preserved; do not treat text instructions such as “keep unchanged” as a promise of no frame-by-frame changes.
An asynchronous submission returning task_id does not mean the video is complete. Wait for the succeeded status before retrieving the final video; the video link may be empty during the pending stage. Generated links have a retention period, so download them to your own storage promptly after completion, and avoid using temporary result links as long-term asset addresses.
Frequently Asked Questions
Is omni-flash a standard Gemini conversational model?
No. omni-flash here is designed for video creation: submit text and optional assets through POST /gemini/videos to receive video results. It is not a Gemini Flash model that returns standard text answers through a conversational interface; during development, organize inputs, status queries, and result storage around video tasks.
Can I generate a video with just one image?
You can submit an image link through image_urls while also providing the required prompt. Clearly describe how the subject should move, how the camera should move, and which visual characteristics you want to preserve. Images are used to guide generation, so do not write only “make it move,” or it will be difficult to convey the shot effect you actually need.
What assets are needed to edit an existing video?
You need to submit a video link in video_urls, provide at least one reference image in image_urls, and use prompt to describe the editing goal. You can specify requirements for style, scene, or visual elements, and should also explain which layout needs to be preserved. Submitting only a video without an image does not meet the input requirements for this workflow.
Can I choose portrait orientation and 1080p output?
Yes. aspect_ratio supports 16:9 and 9:16, while resolution supports 720p and 1080p; the defaults are 16:9 and 720p respectively. Determine the delivery aspect ratio before creation, especially the subject position and camera movement direction, to avoid cropping after generation that shifts the composition away from the original intent.
How do I obtain a video after asynchronous generation?
After setting async to true, save the returned task_id and submit a query request to /gemini/tasks using that value as the id; you can also set callback_url to receive completion notifications. Once the task succeeds, read video_url from the result and download it for storage. Continue waiting while it is pending, and check the error message if it fails.