Gemini Omni Flash Video Generator, Edited by Conversation

Describe a scene and Gemini Omni Flash returns a 720p clip at 24 fps, three to ten seconds long, with an audio track the model generates alongside the picture. Then keep talking: every follow-up instruction re-renders the same clip and leaves anything you didn't mention untouched.

Describe the scene, the motion, and the sound…
16:9/3-10s/720p/With audio

Runs on gemini-omni-flash-preview. Independent interface, not affiliated with Google.

A lone figure on a dune ridge at golden hour

turn 2

“same shot, move it to golden hour”

Every clip this site produces comes from a single model: gemini-omni-flash-preview, the video model Google opened to developers on June 30, 2026. No second engine, no silent substitution.

Five Omni Flash prompts to start from

The artwork below was illustrated by hand, not generated: nothing on these tiles is model output. Each caption is a prompt you can paste in, and each one leans on a different documented behavior — reference tagging, simulated physics, the audio track the model writes itself, a follow-up edit, and timecoded action.

A hiker on a rocky ridge in heavy fog

Use <IMAGE_REF_0> as the hiker: she crests a rocky ridge as fog rolls through, low sun behind her, breath visible in the cold.

A car carving tracks across a snow plain under a low sun

A car drifts wide across a frozen lake, spray lifting behind the rear wheels, tracks curling away under a low sun.

A rain-soaked street at night lit by neon signs

Rain-slick street at night. The neon sign above the door reads OPEN LATE. Tires hiss on wet asphalt and rain taps the awnings. No dialogue.

A marble rolling down a ramp into a chain reaction of blocks

A marble runs down a ramp and topples a row of blocks in sequence. Then, next turn: make the blocks glass, keep everything else the same.

A skateboarder airborne against a low sun on an empty road

[0-2s] the skater drops into the bowl. [2-5s] they go airborne against the sun, shadow stretching down the road.

How conversational video editing actually works

Ask for a change in plain English and Omni Flash returns a new cut of the same clip. Turns run through the Interactions API: passing previous_interaction_id carries the video state forward, so nothing is re-uploaded and anything you didn't mention stays as it was. Keep instructions short — Google notes that overly descriptive edit prompts can lead to unintended changes.

Three chat turns refining one video while the scene persistsgenerate the sceneturn 1make it golden hourturn 2keep her, widen shotturn 3SCENE MEMORY — PERSISTS ACROSS TURNS

Text to video, image to video, or reference to video

Four task modes cover the ways in: text_to_video, image_to_video, reference_to_video, and edit. Tag an upload <FIRST_FRAME> to open the clip on it, or <IMAGE_REF_0> to keep a character or product consistent without pinning the opening shot. Audio uploads are the exception: the API doesn't accept them yet, so direct sound in the prompt instead.

Text, image, video and audio inputs combining into one generated clipTEXTIMAGEVIDEOAUDIOONE CLIP + AUDIO

Does Gemini Omni Flash understand physics?

Google describes an intuitive grasp of gravity, kinetic energy and fluid dynamics, layered onto Gemini's knowledge of history, science and cultural context. Prompts that lean on that get the most out of the model: cloth settling after a gust, sand shifting underfoot, a bottle rocking before it tips. The model card still names complex motion as a known weakness.

A motion arc sampled into frames above a ground planeGRAVITY · MOMENTUM · CONTACTframe 06

Billed by the second, including every edit turn

Google bills 5,792 tokens per second of 720p video at $17.50 per million output tokens — $0.101 a second, or roughly $1.01 for a ten-second clip, the same rate Google charges for Veo 3.1 Fast at 720p. Every conversational edit renders the clip again, so a refined shot costs the first generation plus each turn.

Relative turnaround time compared with slower alternativesTIME TO A USABLE TAKEOMNI FLASHLARGER MODELMANUAL EDIT PASS

What it does, and where it stops

Multi-turn editing

Chain instructions with previous_interaction_id and each turn edits the last result. There are no negative prompts here; unwanted elements come out by asking for the change in the next turn.

Audio in every clip

Each generation returns a synchronized audio track. Steer music, ambience or dialogue in the prompt. Audio files cannot be uploaded as references, and voice editing is unsupported.

720p at 24 fps

Output is 720p, 24 fps MP4, three to ten seconds long. That is the only resolution Google documents today, so treat any 1080p or 4K claim as false.

Two aspect ratios

The aspect_ratio parameter accepts 16:9 and 9:16, nothing else. Landscape is the default; portrait covers Shorts, Reels and TikTok. There is no 1:1, 4:3 or 4:5 option.

Reference image tags

Tag an upload <FIRST_FRAME> to start the clip on it, or <IMAGE_REF_0> to carry a face, product or location through the shot. Google's enterprise docs put the ceiling at ten images per prompt.

Text inside the frame

Ask for signage, captions or a title card and the model renders text Google describes as legible. Its model card is equally clear that perfectly accurate text is still unsolved.

Timing and shot control

Write events against timecodes, in brackets or plain words. By default the model cuts between several shots, so say "a single unbroken take" when you want one.

Always-on SynthID

Every generated clip carries an imperceptible SynthID watermark that can be detected programmatically. There is no documented opt-out, so treat "watermark-free Omni Flash output" as a false claim.

Regional editing limits

Editing uploaded video is unavailable in the EEA, Switzerland and the UK; editing clips the model generated still works there. Video extension and frame interpolation are unsupported everywhere.

How to use it in three steps

  1. 1

    Write the prompt and tag your references

    Describe the scene, the motion, and the sound you want. If you're starting from stills, tag each upload so the model knows its job: one image as the first frame, the rest as visual references. Set the mode explicitly — text, image, reference, or edit to video — or let the prompt imply it.

  2. 2

    Set the frame, the timing, and the audio

    Pick 16:9 or 9:16; those are the only two ratios the model accepts. Clips run three to ten seconds. By default it cuts between several shots, so ask for a single continuous take if you want one. Time events in the prompt, like [0-3s] and [3-6s]. Audio is generated with every clip, so say what it should sound like.

  3. 3

    Refine it by chatting, then download

    Watch the result, then type the one thing you want changed. The next turn reuses the stored video state, so unmentioned elements stay put; adding "keep everything else the same" helps. Short instructions beat long ones. Each turn renders a new clip and spends credits for its full length. Export the finished MP4 when it lands.

The OmniFlash studio: prompt panel, preview canvas and timelinePROMPTREFERENCESOUTPUT16:98s720pGeneratePREVIEWTIMELINE0s2s4s6s8s

Gemini Omni Flash vs Veo 3.1

Every cell is sourced from Google's published documentation or marked as not documented. We don't print specs we can't trace to a primary source. Verified against Google's published documentation for gemini-omni-flash-preview on 2026-08-08.

DimensionGemini Omni FlashVeo 3.1
Output resolution720p at 24 fps, the only documented optionPriced for 720p, 1080p and 4K output
Clip length3 to 10 seconds per generationNot stated on the Google pages we cite
Native audioGenerated with every clip, no separate audio chargeGoogle prices Veo 3.1 tiers "with audio"
Conversational multi-turn editingYes, chained with previous_interaction_id; Google's launch post describes stacking up to three sequential editsGoogle points to Veo 3.1 for scene extension and last-frame control instead
Published price per second, 720p with audio$0.1014 (5,792 tokens/second at $17.50 per 1M video output tokens)720p with audio: $0.40 Standard, $0.10 Fast, $0.05 Lite (higher rates at 1080p and 4K)

Where it earns its place in a workflow

Three to ten seconds, 720p, audio included, and edits you make by talking. That shape suits some jobs and not others, so here is where it actually pays off and what to watch for in each case.

Vertical clips for Shorts, Reels and TikTok

9:16 is a first-class output here, not a crop, so a portrait clip comes out framed for the feed. Ten seconds is the ceiling, which suits a hook, a product beat, or a transition rather than a full edit. Audio arrives with the clip, and you can redirect the music or drop dialogue in a follow-up turn.

A four-shot sequence with direction notes under each frameSHOT SEQUENCEWIDEMIDCLOSEPUSHtimecoded beats · consistent characters

Ad and product variants without a round trip to an editor

Feed in a product still as the first frame, then ask for the variant you need: different lighting, a different background, different wording on the sign. Each request is one instruction, and the parts you didn't mention carry over. Text rendering is legible rather than typeset, so proof every frame before a client sees it.

One product shot re-lit and re-staged as three campaign variantsONE SOURCE → THREE VARIANTSABCrelightnew backdropedit signagebrand consistency held across every edit

Developers sizing up the gemini-omni-flash-preview API

The model runs through the Interactions API, and turns chain with previous_interaction_id rather than re-uploading video. Worth knowing before you budget: system instructions, temperature and negative prompts are unavailable, video comes back as inline base64 unless you ask for delivery as a URI (Google recommends that above about 4MB), and every edit turn bills as a full regeneration.

Vertical nine-by-sixteen frames with an audio waveform beneath9:16 · BUILT FOR FEEDS16:9 TOOnative audio generated with the clip

Gemini Omni Flash Pricing, Billed by the Second

One credit buys one second of 720p video with audio. A five-second clip spends five credits, a ten-second clip spends ten, and each conversational edit renders the clip again, so it costs its full length a second time. Credits reset monthly.

Starter

Enough runway to find out whether the model fits your work.

$13/month

40 credits a month = 40 seconds of 720p video, about 8 five-second clips or 4 ten-second clips.

  • gemini-omni-flash-preview, 720p at 24 fps, clips of 3 to 10 seconds
  • Text to video and image to video, in 16:9 or 9:16
  • Conversational edit turns, each charged at the clip's length in credits
  • MP4 download, with Google's SynthID watermark on every file
  • Credits reset each month and do not roll over

Creator

For a steady weekly output of short clips.

$30/month

100 credits a month = 100 seconds of 720p video, about 20 five-second clips or 10 ten-second clips.

  • Everything in Starter
  • All four modes: text, image, reference, and edit to video
  • Reference images tagged as first frame or visual reference
  • 3 generations running at once
  • Runs on Google's paid tier, which Google does not use to improve its products
  • Top-up credits available when a month runs short

Studio

For teams shipping variant sets every week.

$50/month

170 credits a month = 170 seconds of 720p video, about 34 five-second clips or 17 ten-second clips.

  • Everything in Creator
  • Edit footage you upload yourself, where Google's regional rules allow it
  • 6 generations running at once
  • Generations kept private to your account
  • Per-clip usage log showing seconds and credits spent

Production

Team seats, shared credits, and usage you can export.

$250/month

850 credits a month = 850 seconds of 720p video, about 170 five-second clips or 85 ten-second clips.

  • Everything in Studio
  • 12 generations running at once
  • Monthly usage export with seconds and cost per clip
  • Seats for your team on one credit pool
  • Invoice billing on request

Plans start at $13 a month. Google's own API has no free tier for this model.

Gemini Omni Flash questions, answered

Verified against Google's model page, API guide and pricing table on 2026-08-08. Where Google hasn't published something, we say so rather than guess.

Not through the API. Google's pricing table lists the free tier as "Not available" for gemini-omni-flash-preview, so every generated second carries a real cost, roughly $0.1014 per second of 720p output. Google's launch post says the model is rolling out at no cost inside YouTube Shorts and the YouTube Create app; the Gemini app and Google Flow need a Google AI subscription.

One model, six ways in

Everything here runs on gemini-omni-flash-preview. No second engine behind the buttons.

Text to video
Image to video
Reference to video
Conversational edit turns
16:9 and 9:16 output
Native audio on every clip

Write one prompt, then change your mind out loud

720p at 24 fps, three to ten seconds, landscape or portrait, audio generated with the picture. Refine by replying instead of starting the shot over.