Gemini Omni Flash Video Generator, Edited by Conversation
Describe a scene and Gemini Omni Flash returns a 720p clip at 24 fps, three to ten seconds long, with an audio track the model generates alongside the picture. Then keep talking: every follow-up instruction re-renders the same clip and leaves anything you didn't mention untouched.
Runs on gemini-omni-flash-preview. Independent interface, not affiliated with Google.
turn 2
“same shot, move it to golden hour”
Every clip this site produces comes from a single model: gemini-omni-flash-preview, the video model Google opened to developers on June 30, 2026. No second engine, no silent substitution.
Five Omni Flash prompts to start from
The artwork below was illustrated by hand, not generated: nothing on these tiles is model output. Each caption is a prompt you can paste in, and each one leans on a different documented behavior — reference tagging, simulated physics, the audio track the model writes itself, a follow-up edit, and timecoded action.
How conversational video editing actually works
Ask for a change in plain English and Omni Flash returns a new cut of the same clip. Turns run through the Interactions API: passing previous_interaction_id carries the video state forward, so nothing is re-uploaded and anything you didn't mention stays as it was. Keep instructions short — Google notes that overly descriptive edit prompts can lead to unintended changes.
Text to video, image to video, or reference to video
Four task modes cover the ways in: text_to_video, image_to_video, reference_to_video, and edit. Tag an upload <FIRST_FRAME> to open the clip on it, or <IMAGE_REF_0> to keep a character or product consistent without pinning the opening shot. Audio uploads are the exception: the API doesn't accept them yet, so direct sound in the prompt instead.
Does Gemini Omni Flash understand physics?
Google describes an intuitive grasp of gravity, kinetic energy and fluid dynamics, layered onto Gemini's knowledge of history, science and cultural context. Prompts that lean on that get the most out of the model: cloth settling after a gust, sand shifting underfoot, a bottle rocking before it tips. The model card still names complex motion as a known weakness.
Billed by the second, including every edit turn
Google bills 5,792 tokens per second of 720p video at $17.50 per million output tokens — $0.101 a second, or roughly $1.01 for a ten-second clip, the same rate Google charges for Veo 3.1 Fast at 720p. Every conversational edit renders the clip again, so a refined shot costs the first generation plus each turn.
What it does, and where it stops
Multi-turn editing
Chain instructions with previous_interaction_id and each turn edits the last result. There are no negative prompts here; unwanted elements come out by asking for the change in the next turn.
Audio in every clip
Each generation returns a synchronized audio track. Steer music, ambience or dialogue in the prompt. Audio files cannot be uploaded as references, and voice editing is unsupported.
720p at 24 fps
Output is 720p, 24 fps MP4, three to ten seconds long. That is the only resolution Google documents today, so treat any 1080p or 4K claim as false.
Two aspect ratios
The aspect_ratio parameter accepts 16:9 and 9:16, nothing else. Landscape is the default; portrait covers Shorts, Reels and TikTok. There is no 1:1, 4:3 or 4:5 option.
Reference image tags
Tag an upload <FIRST_FRAME> to start the clip on it, or <IMAGE_REF_0> to carry a face, product or location through the shot. Google's enterprise docs put the ceiling at ten images per prompt.
Text inside the frame
Ask for signage, captions or a title card and the model renders text Google describes as legible. Its model card is equally clear that perfectly accurate text is still unsolved.
Timing and shot control
Write events against timecodes, in brackets or plain words. By default the model cuts between several shots, so say "a single unbroken take" when you want one.
Always-on SynthID
Every generated clip carries an imperceptible SynthID watermark that can be detected programmatically. There is no documented opt-out, so treat "watermark-free Omni Flash output" as a false claim.
Regional editing limits
Editing uploaded video is unavailable in the EEA, Switzerland and the UK; editing clips the model generated still works there. Video extension and frame interpolation are unsupported everywhere.
How to use it in three steps
- 1
Write the prompt and tag your references
Describe the scene, the motion, and the sound you want. If you're starting from stills, tag each upload so the model knows its job: one image as the first frame, the rest as visual references. Set the mode explicitly — text, image, reference, or edit to video — or let the prompt imply it.
- 2
Set the frame, the timing, and the audio
Pick 16:9 or 9:16; those are the only two ratios the model accepts. Clips run three to ten seconds. By default it cuts between several shots, so ask for a single continuous take if you want one. Time events in the prompt, like [0-3s] and [3-6s]. Audio is generated with every clip, so say what it should sound like.
- 3
Refine it by chatting, then download
Watch the result, then type the one thing you want changed. The next turn reuses the stored video state, so unmentioned elements stay put; adding "keep everything else the same" helps. Short instructions beat long ones. Each turn renders a new clip and spends credits for its full length. Export the finished MP4 when it lands.
Gemini Omni Flash vs Veo 3.1
Every cell is sourced from Google's published documentation or marked as not documented. We don't print specs we can't trace to a primary source. Verified against Google's published documentation for gemini-omni-flash-preview on 2026-08-08.
| Dimension | Gemini Omni Flash | Veo 3.1 |
|---|---|---|
| Output resolution | 720p at 24 fps, the only documented option | Priced for 720p, 1080p and 4K output |
| Clip length | 3 to 10 seconds per generation | Not stated on the Google pages we cite |
| Native audio | Generated with every clip, no separate audio charge | Google prices Veo 3.1 tiers "with audio" |
| Conversational multi-turn editing | Yes, chained with previous_interaction_id; Google's launch post describes stacking up to three sequential edits | Google points to Veo 3.1 for scene extension and last-frame control instead |
| Published price per second, 720p with audio | $0.1014 (5,792 tokens/second at $17.50 per 1M video output tokens) | 720p with audio: $0.40 Standard, $0.10 Fast, $0.05 Lite (higher rates at 1080p and 4K) |
Where it earns its place in a workflow
Three to ten seconds, 720p, audio included, and edits you make by talking. That shape suits some jobs and not others, so here is where it actually pays off and what to watch for in each case.
Vertical clips for Shorts, Reels and TikTok
9:16 is a first-class output here, not a crop, so a portrait clip comes out framed for the feed. Ten seconds is the ceiling, which suits a hook, a product beat, or a transition rather than a full edit. Audio arrives with the clip, and you can redirect the music or drop dialogue in a follow-up turn.
Ad and product variants without a round trip to an editor
Feed in a product still as the first frame, then ask for the variant you need: different lighting, a different background, different wording on the sign. Each request is one instruction, and the parts you didn't mention carry over. Text rendering is legible rather than typeset, so proof every frame before a client sees it.
Developers sizing up the gemini-omni-flash-preview API
The model runs through the Interactions API, and turns chain with previous_interaction_id rather than re-uploading video. Worth knowing before you budget: system instructions, temperature and negative prompts are unavailable, video comes back as inline base64 unless you ask for delivery as a URI (Google recommends that above about 4MB), and every edit turn bills as a full regeneration.
Gemini Omni Flash Pricing, Billed by the Second
One credit buys one second of 720p video with audio. A five-second clip spends five credits, a ten-second clip spends ten, and each conversational edit renders the clip again, so it costs its full length a second time. Credits reset monthly.
Starter
Enough runway to find out whether the model fits your work.
40 credits a month = 40 seconds of 720p video, about 8 five-second clips or 4 ten-second clips.
- gemini-omni-flash-preview, 720p at 24 fps, clips of 3 to 10 seconds
- Text to video and image to video, in 16:9 or 9:16
- Conversational edit turns, each charged at the clip's length in credits
- MP4 download, with Google's SynthID watermark on every file
- Credits reset each month and do not roll over
Creator
For a steady weekly output of short clips.
100 credits a month = 100 seconds of 720p video, about 20 five-second clips or 10 ten-second clips.
- Everything in Starter
- All four modes: text, image, reference, and edit to video
- Reference images tagged as first frame or visual reference
- 3 generations running at once
- Runs on Google's paid tier, which Google does not use to improve its products
- Top-up credits available when a month runs short
Studio
For teams shipping variant sets every week.
170 credits a month = 170 seconds of 720p video, about 34 five-second clips or 17 ten-second clips.
- Everything in Creator
- Edit footage you upload yourself, where Google's regional rules allow it
- 6 generations running at once
- Generations kept private to your account
- Per-clip usage log showing seconds and credits spent
Production
Team seats, shared credits, and usage you can export.
850 credits a month = 850 seconds of 720p video, about 170 five-second clips or 85 ten-second clips.
- Everything in Studio
- 12 generations running at once
- Monthly usage export with seconds and cost per clip
- Seats for your team on one credit pool
- Invoice billing on request
Plans start at $13 a month. Google's own API has no free tier for this model.
Gemini Omni Flash questions, answered
Verified against Google's model page, API guide and pricing table on 2026-08-08. Where Google hasn't published something, we say so rather than guess.
More resources
What a Gemini Omni Flash Clip Really Costs
The token math behind $0.1014 per second, then the number nobody publishes: what one usable clip costs after re-rolls and edit turns.
Prompting Gemini Omni Flash: Timecodes, Reference Tags and 9:16
How to force a single continuous shot, time events inside the prompt, tag first frames versus references, and direct the generated audio track.
Every Documented Gemini Omni Flash Limit
No video extension, no interpolation, no audio upload, no voice editing, no negative prompts, plus the EEA, Swiss and UK editing carve-out.
One model, six ways in
Everything here runs on gemini-omni-flash-preview. No second engine behind the buttons.
Write one prompt, then change your mind out loud
720p at 24 fps, three to ten seconds, landscape or portrait, audio generated with the picture. Refine by replying instead of starting the shot over.