Gemini Omni Flash is Google's unified video model for generating and editing clips with native sound. One model handles text-to-video, image-to-video, reference-to-video, and video editing, including follow-up edits to an earlier Omni generation. It is now available across Creativly's Video sessions, Agent, and Flow canvas.
What is Gemini Omni Flash?
Gemini Omni Flash is Google's unified AI video model, now built into Creativly. Instead of juggling separate tools for generating, animating, referencing, and editing video, you pick one of four video tasks and stay on the same model from first draft to final edit.
The practical benefit is continuity. You can start with a prompt, animate a still, direct a scene with reference images, or edit a clip — all without switching models. And when you edit a clip Gemini Omni Flash just made, it picks up where the last generation left off instead of starting over from scratch.
In Creativly, Gemini Omni Flash handles four kinds of video from one place — generate from a text prompt, animate a still image, guide a scene with reference images, or edit an existing clip — all with native sound, and without switching between different models.
One model takes an idea from first frame to finished edit, while Creativly keeps every job running in the background and visible in the same project.
The four supported Gemini Omni video tasks
1. Text-to-video
Start with a written direction and no media input. This is the cleanest route for concept shots, establishing scenes, motion studies, and rapid visual exploration. Describe the subject, action, camera, environment, lighting, and sound in one prompt.
2. Image-to-video
Supply one image as the starting visual and describe how it should move. This works well for product photography, character art, campaign stills, and any approved frame that needs motion without rebuilding the composition from text alone.
3. Reference-to-video
Provide one or more reference images to guide the generated scene. Gemini Omni accepts up to 10 images in total, making this task useful when a prompt needs stronger visual grounding for a person, product, wardrobe, environment, or art direction. Reference images guide a new generation; they are not first-and-last-frame interpolation.
4. Video editing
Upload one video and describe the change, or continue from a compatible video previously generated with Gemini Omni. The edit task is suited to directed visual revisions such as changing atmosphere, restyling a scene, adjusting visible details, or refining motion with another instruction.
When you edit a clip you just generated, Creativly keeps the edit connected to your original so follow-ups build on the right source — and only ever on your own work. You can also upload any other compatible video and edit that directly.
Native audio is part of the output
Gemini Omni generates sound with the video instead of returning a silent clip that always needs a separate audio pass. Put the sound direction in the same prompt as the visual direction: room tone, weather, footsteps, machinery, traffic, or another environmental cue that belongs in the scene.
Native audio means the soundtrack is generated together with the video — it does not take an uploaded soundtrack or voice clip as a conditioning input, and Creativly does not expose unsupported audio-reference controls for this model.
Gemini Omni Flash limits
Gemini Omni Flash supports a focused set of output options:
- Duration: 3 to 10 seconds, selected in whole-second steps.
- Resolution: 720p.
- Aspect ratio: 16:9 landscape or 9:16 vertical.
- Image inputs: up to 10 images for reference-guided generation; image-to-video uses one primary image.
- Video inputs: one uploaded video for editing.
- Audio: generated natively in the output; uploaded audio references are not supported.
The model does not currently expose 1080p or 4K output, square video, a negative prompt, last-frame interpolation, video-reference motion guidance, uploaded audio conditioning, or video extension controls. Creativly validates these limits before submission instead of silently dropping extra media or coercing an unsupported setting.
How to use Gemini Omni in Creativly
Gemini Omni Flash is wired into the same three creation surfaces as the other video models:
- Video sessions: choose Gemini Omni Flash, set duration and aspect ratio, attach the media required by the task, and submit. Jobs update asynchronously in the session while you continue working.
- Agent: describe the video or revision in chat. It remembers which model and which clip you're working from, so an approved follow-up edits the right output.
- Flow: use a video node for repeatable graphs. Connect a prompt, one starting image, reference images, or one editable video through the existing media handles, then branch the result into the next node.
Practical Gemini Omni workflows
Campaign still to vertical motion
Generate or upload an approved campaign still, pass it to Gemini Omni as the starting image, choose 9:16, and prompt the subject motion, camera movement, and environmental sound. Keep the original image node in Flow so the source remains reusable for other cuts.
Reference pack to product scene
Attach product angles, packaging, wardrobe, and location references, then generate a new 16:9 scene. Use only the images that materially define the shot; a smaller, coherent reference set usually gives clearer direction than filling all 10 slots.
Generate, review, then revise in Agent
Ask Agent for a first clip, review the result, and request a focused change in the same session. Creativly keeps the edit tied to your original clip, so the revision always builds on the right source.
Uploaded clip to controlled restyle
Start from one uploaded clip, choose edit mode, and specify what should change and what must stay the same. This works for any compatible video — including clips made with a different model — not just ones generated with Gemini Omni.
Prompting tips for Gemini Omni video
- Write the action as a sequence. State what moves first, how the camera responds, and how the shot ends.
- Separate visual and sound direction. A short final sentence for ambience or effects makes the audio intent easier to parse.
- For edits, define invariants. Say what must remain unchanged before describing the revision.
- Match scope to duration. A 3-second shot should contain one clear beat; use 8 to 10 seconds for a longer movement or reveal.
- Use references with distinct jobs. Choose one image for the subject, another for product detail, and another for environment instead of near-duplicate frames.
Gemini Omni compared with other video models
Gemini Omni's differentiator is not that it replaces every specialist model. It is the combination of four video tasks, native audio, and conversational editing in one place. Models such as Veo may be a better fit when a workflow needs a different resolution or generation profile, while Seedance 2.0 supports a broader set of multimodal reference inputs.
Flow makes that tradeoff practical: keep the prompt and source assets in the graph, run the model that matches the shot, and compare results without rebuilding the workflow.
Gemini Omni Flash pricing
Creativly prices Gemini Omni by requested output duration, so a longer clip uses more platform credits than a shorter one. BYOK runs use the connected Gemini credential instead of platform generation credits. See the pricing page for the current rate and plan details.
Treat Gemini Omni Flash as a production-ready video model in your workflow: review outputs before publishing and keep source assets in the project so you can iterate and re-run quickly.
FAQ
Who makes Gemini Omni Flash?
Google. It is part of the Gemini model family, with text-to-video, image-to-video, reference-guided generation, and video editing in one model — all with native sound.
Can Gemini Omni edit a video?
Yes. You can edit one uploaded video with a text instruction, or keep editing a clip you just generated with Gemini Omni — each revision builds on the last.
Does Gemini Omni generate audio?
Yes. Audio is generated natively with the video. Gemini Omni Flash does not accept an uploaded audio clip as a reference input.
What video length and resolution does it support?
Outputs are 3 to 10 seconds at 720p, in either 16:9 landscape or 9:16 vertical format.
How many reference images can I use?
Up to 10 image inputs. Image-to-video uses one primary image; reference-to-video can use a set of images to guide a newly generated scene.
Where is Gemini Omni available in Creativly?
In Video sessions, the Creativly Agent, and the Flow node canvas.
Try Gemini Omni Flash
Open a Video session for a direct generation, use Agent for conversational revisions, or build a reusable graph in Flow.



