Gemini Omni Flash: What It Is, What It Costs, How It Compares
Gemini Omni Flash is Google DeepMind's video generation and editing model. In the Studio it makes 3 to 10 second clips with sound at 360p to 4K, from 5 credits. Checked against Google's documentation and the model's API schema on 3 October 2026.

Gemini Omni Flash is Google DeepMind's video model. It takes text, images, audio and video as input and returns video with a soundtrack. In the Studio it runs as version 1.1, making clips of 3 to 10 seconds at 360p, 720p, 1080p or 4K, from 5 credits for 3 seconds at 360p. The default clip, 8 seconds at 720p, costs 48 credits.
That is the answer most people searching for it want. The rest of this post is what Google says the model is for, what each setting costs, and where it sits next to Veo 3.1 and Seedance 2.0.
There was no page on this site for it until today, although people were already searching for one. Our search data logged ten distinct questions about it. This post answers the ones the facts support.
The Short Version
- What it is: Google's model for generating and editing video from any mix of text, image, audio and video.
- Length: 3 to 10 seconds, in whole seconds.
- Resolution: 360p, 720p, 1080p or 4K. Google's own API documentation labels the 1080p and 4K options "upscaled".
- Sound: always generated. There is no switch to turn it off.
- Price: 5 to 180 credits a clip, depending on length and resolution.
What Google Says It Is
Google DeepMind's model card, first published in May 2026 and updated in August, calls Gemini Omni Flash "our next step towards models that can create and edit anything from any input", starting with video. It lists the inputs as text, images, audio and video files, and the output as "high-quality, high-resolution video with audio".
The developer timeline, from Google's Gemini API changelog:
- 30 June 2026:
gemini-omni-flash-previewreleased, "a high-performance multimodal model designed for high-speed video generation and conversational video editing". - 27 August 2026:
gemini-omni-1.1-flashreleased, "the GA version of our fast, conversational video generation and editing model". The same entry notes that "1080p and 4K outputs are generated using upscaling", and that the preview endpoint would be deprecated on 30 September 2026.
Google's announcement of 1.1, dated 27 August, lists what changed: scene extension in 10-second steps up to 40 seconds in total, first and last frame control, video references of up to three seconds, 360p drafts, and 1080p or 4K output.
The model card is also candid about the weak spots. In Google's words, "maintaining complete consistency throughout edits, generating scenes with complex motion, or rendering perfectly accurate text remains a challenge." That is the maker talking. Read it before you build a brief around on-screen lettering.
Two more details from Google's documentation are worth knowing before you spend anything. Google says every generated video carries a SynthID watermark, invisible to viewers but detectable by software. And English is the only prompt language Google lists as fully supported; others "may work but results can vary".
What the Studio Runs
Not everything in Google's announcement reaches the Studio. Conversational editing and scene extension, both offered through Google's own API, are not in it. What the Studio does run is three modes, checked against the model's API schema on 3 October 2026:
| Gemini Omni Flash 1.1 in the Studio | |
|---|---|
| Modes | Text, Image (start frame, optional end frame), Reference |
| Length | 3 to 10 seconds, any whole number, default 8 |
| Resolution | 360p, 720p (default), 1080p, 4K |
| Aspect ratio | 16:9 or 9:16, in every mode |
| Sound | Always on, no switch |
| References | Up to 10 images and up to 3 video clips |
A few rows deserve a note.
Sound is not optional. The endpoint has no field to request or suppress audio, so the Studio shows no toggle. Google's prompt guide says the model "will try to generate an appropriate audio track" by default, and that you steer it by describing the sound you want in the prompt. If the clip is going under your own music, plan to strip the track in the edit.
Reference clips are short. Google's documentation allows a maximum of three clips, "up to 3 seconds each", and says any audio in a reference clip is ignored.
Two shapes only. No square, no 4:3, no 21:9. For a feed post that needs 1:1, pick another model before you write the prompt.
What Each Clip Costs
Every figure below is the live Studio price in credits, shown in the console before you press Generate.
| Resolution | 3 seconds | 5 seconds | 8 seconds | 10 seconds |
|---|---|---|---|---|
| 360p | 5 | 9 | 14 | 18 |
| 720p | 18 | 30 | 48 | 60 |
| 1080p | 27 | 45 | 72 | 90 |
| 4K | 54 | 90 | 144 | 180 |
Three things follow from that table.
- 360p is the sketchpad. A 5-second test costs 9 credits, under a third of the same clip at 720p. Google pitches the 360p tier the same way, as a cheaper drafting resolution.
- A higher resolution is a new take. In the Studio, choosing 1080p or 4K starts a fresh generation. It does not upscale the 360p draft you liked, so the sharper clip will not be the same clip. Use 360p to settle the prompt, not the performance.
- 4K costs ten times 360p. 180 credits for 10 seconds, against 18. Since Google itself calls its 4K output upscaled, spend that only where the delivery spec asks for it.
For what the model costs the provider per second, set against every other model on the roster, see what AI video costs per second. For what a finished shot costs once retakes are counted, the Render Ledger does the arithmetic.
Is It Free?
Not on Google's developer API. Its pricing page, checked on 3 October 2026, lists the free tier for gemini-omni-1.1-flash as "Not available".
In the Studio, a free account refills to 9 credits a day. That covers one short clip: up to 3 seconds of Wan 3.0 at 480p, or 5 seconds of Gemini Omni Flash at 360p. It is enough to see whether the model suits your idea. It is not enough for a deliverable, and anything longer or sharper costs more than the free daily allowance.
Gemini Omni Flash vs Veo 3.1
Both are Google models, and both generate sound with the picture. In the Studio they differ mainly in length, controls and price:
| Gemini Omni Flash 1.1 | Veo 3.1 | |
|---|---|---|
| Length | 3 to 10 seconds | 4, 6 or 8 seconds |
| Resolution | 360p to 4K | 720p, 1080p, 4K |
| Sound | Always on | On or off |
| Reference inputs | 10 images, 3 clips | 3 images |
| 8 seconds at 720p | 48 credits | 200 credits |
The price gap is the obvious line. Veo 3.1 costs 25 credits a second at 720p and Omni costs 6, so the same 8-second clip is 200 credits against 48. Veo 3.1 Fast sits between them at 80 credits for 8 seconds, and Veo 3.1 Lite at 720p comes in at 20.
What the table does not tell you is which one looks better on your shot. We have not run the two side by side on a controlled prompt set, and we are not going to guess. Our Veo 3 review covers what the Veo line does well, dialogue in particular. Omni's case rests on Google's description of it: a model built to take mixed inputs, which in the Studio means reference clips as well as stills.
So the practical split is about inputs and length rather than quality. If the shot needs a 10-second take, or reference clips alongside stills, Omni is the one of the two that does it in the Studio. If you need a clip without a soundtrack, Veo 3.1 has the switch.
Gemini Omni Flash vs Seedance 2.0
| Gemini Omni Flash 1.1 | Seedance 2.0 | |
|---|---|---|
| Length | 3 to 10 seconds | 4 to 15 seconds |
| Resolution | 360p to 4K | 480p to 4K |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 |
| Sound | Always on | On or off |
| Reference inputs | 10 images, 3 clips | 9 images |
| 5 seconds at 720p | 30 credits | 91 credits |
Seedance 2.0 wins on reach. It runs half as long again, and it is the one to pick for a square post or a 21:9 frame. Omni wins on price, at roughly a third of the cost for the same 5 seconds at 720p, and it is the only one of the two that accepts reference clips.
Again, that is a spec and price comparison, not a verdict on output. Seedance 2.0 is covered alongside its successor in our Seedance 2.5 review. If you need a long take with sound for very little, Wan 3.0 runs to 30 seconds and is worth a look too.
Gemini 3.5 Flash vs Gemini Omni
People search for this because the names share a word. Google's own documentation settles it, and it is not a contest.
Gemini 3.5 Flash is a language model. Google's model page lists its inputs as text, image, video, audio and PDF, and its output as text; it lists image generation and audio generation as not supported. It reads and reasons. It does not make video.
Gemini Omni Flash is a video model. Google's model page for gemini-omni-1.1-flash lists its output as video, 3 to 10 seconds, at 360p, 720p, 1080p or 4K and 24 frames a second.
One is not a newer or cheaper version of the other. They are different tools that happen to share the Gemini name. Gemini 3.5 Flash is not in the Studio, because the Studio makes images and video.
How to Brief It
A caveat first. We have no observed prompting notes of our own for Omni yet. The Studio's Prompt Director carries published provider guidance for it, labelled as such rather than as something we have confirmed. So treat what follows as Google's advice, passed on.
Ask for one shot if you want one shot. Google's guide warns that "by default Omni Flash will try to create a video with a few different shots." For a single take, say so: "in a single continuous shot", "no scene cuts".
Describe the sound. Because audio always comes back, an unscripted soundtrack is the default. Name the music, the ambience, or "no dialogue".
Time events in plain language. Google's examples include "After 3 seconds, a woman enters the scene" and a timecode style such as [0-3s], [3-6s], [6-10s].
A brief built that way looks like this:
In a single continuous shot, a ceramicist in a sunlit workshop
lifts a finished bowl from the wheel and turns it towards the
camera. Slow push in, shallow depth of field. Sound: the wheel
winding down, a radio playing softly in the background. No dialogue.
What we would pick: Text mode, 360p and 8 seconds while the prompt settles (14 credits), then 1080p for the take you keep (72 credits).
What We Have Not Tested
Plainly: we have checked the API schema, Google's documentation, the list price and the Studio's credit ladder. We have not run a controlled comparison of Omni against Veo 3.1, Seedance 2.0 or anything else, and we have no measured first-try hit rate for it. Nothing above is a quality verdict.
The open questions we most want answered:
- How the 1080p and 4K tiers, which Google labels upscaled, compare in fine detail with a model that renders those sizes directly.
- How well the always-on audio follows a prompt that asks for silence or a specific piece of music.
- Whether Reference mode holds a face across shots as well as Google's examples suggest.
When we have answers, they go here, dated.
Questions About Gemini Omni Flash
What is Gemini Omni Flash? Gemini Omni Flash is Google DeepMind's video generation and editing model. It takes text, images, audio and video as input and returns video with sound. Google released the generally available version, gemini-omni-1.1-flash, on 27 August 2026, and the AIVideos Studio runs version 1.1 for clips of 3 to 10 seconds at 360p, 720p, 1080p or 4K.
How much does Gemini Omni Flash cost? In the AIVideos Studio, Gemini Omni Flash costs from 5 credits for 3 seconds at 360p. The default clip, 8 seconds at 720p, costs 48 credits, and the top setting, 10 seconds at 4K, costs 180 credits. The price is shown before you generate, and credits never expire.
Is Gemini Omni Flash free? Not on Google's developer API, whose pricing page lists no free tier for it. In the AIVideos Studio, the free daily credits buy one short clip: up to 3 seconds of Wan 3.0 at 480p, or 5 seconds of Gemini Omni Flash at 360p. Anything longer or sharper costs more than the free daily allowance.
Gemini Omni Flash vs Veo 3.1: what is the difference? Both are Google models that generate sound with the picture. In the Studio, Gemini Omni Flash runs 3 to 10 seconds and costs 48 credits for 8 seconds at 720p, while Veo 3.1 runs 4, 6 or 8 seconds and costs 200 credits for the same clip. Veo 3.1 has an audio switch and Omni does not. We have not compared their output on a controlled prompt set.
Gemini Omni Flash vs Seedance 2.0: what is the difference? Seedance 2.0 reaches 15 seconds and six aspect ratios, including 21:9, and takes up to 9 reference images. Gemini Omni Flash stops at 10 seconds and offers 16:9 or 9:16, but takes up to 10 reference images plus 3 reference clips. In the Studio a 5-second clip at 720p costs 30 credits on Omni and 91 on Seedance 2.0.
What is the difference between Gemini 3.5 Flash and Gemini Omni? They do different jobs. Google's documentation describes Gemini 3.5 Flash as a model that reads text, images, video, audio and PDFs and answers in text. Gemini Omni Flash is Google's video generation and editing model, and its output is video. One is not a cheaper or newer version of the other, so there is no like-for-like comparison to make.
Sources
All checked on 3 October 2026.
- Google DeepMind, Gemini Omni Flash model card
- Google, Gemini API changelog and Gemini API pricing
- Google, Gemini Omni Flash model page and video generation guide
- Google, Gemini 3.5 Flash model page
- Google, Gemini Omni 1.1 Flash lets you build with more control, 27 August 2026
- The model's API schema and published rate on fal.ai, the public price source behind our Render Ledger
Where to Start
Open the Studio, pick Gemini Omni Flash, set it to 360p and write the shot. At 9 credits for 5 seconds it is one of the cheapest ways on the roster to see an idea move. If you would rather hand the brief to someone who runs these models every day, that is what our production services are for.
Generate cinematic AI video, from €15
31 AI models. No subscription. Buy credits, generate on demand, own the results outright.
