
Gemini Omni 1.1 Flash vs Veo 3.1 API: Pricing, Benchmarks, and Which to Use
Gemini Omni 1.1 Flash leads Veo 3.1 by ~150 Elo on Arena text-to-video. Pricing, specs, and which Google video API fits your pipeline in 2026.
Google now sells two current video models, and picking between them is a real decision. Gemini Omni 1.1 Flash launched August 27, 2026 and took #1 on Arena text-to-video a day later. Veo 3.1 has been the workhorse since November 2025: keyframes on every tier, Extend for chained narratives, dedicated upscalers, and a draft tier at $0.075 per video. Both are live on Unifically. Here is the head-to-head, with the numbers as of August 28, 2026.
TL;DR: For most new projects, pick Gemini Omni 1.1 Flash: it leads Veo 3.1's best Arena row by about 150 Elo on text-to-video (1515 vs 1364), takes richer references (7 images plus 3 named characters vs Veo's 3 images), supports start and end keyframes, and it is the only one that edits an existing video. Pick Veo 3.1 when your pipeline chains narratives with the Extend endpoint or lives on the cheapest drafts: Lite Relaxed at $0.075 per video still undercuts Omni's $0.225 360p draft 3 to 1. Omni generation costs $0.45 to $0.90 per video on Unifically; Veo runs $0.075 to $0.60.
Key takeaways
- Arena text-to-video, August 28, 2026: Gemini Omni 1.1 Flash #1 at 1515 Elo; Veo 3.1's best row #10 at 1364. On image-to-video: Omni #2 at 1488, Veo #11 at 1398.
- Veo 3.1 pricing on Unifically is flat per video: $0.075 (Lite Relaxed), $0.15 (Lite), $0.30 (Fast), $0.60 (Quality). Omni 1.1 generation is priced by duration: $0.45 (4s) to $0.90 (10s) with 1080p and 4K at no extra cost, and 360p drafts at half price.
- Omni takes up to 7 reference images plus 3 named character references and has a video-edit endpoint; Veo takes up to 3 reference images (Fast and Lite tiers only) and has no edit mode.
- Both support start and end frame keyframes. Veo has Extend for chaining videos into longer narratives; Omni maxes at 10 seconds per generation.
- Both output video with synchronized audio in one file, in 16:9 or 9:16. Veo runs 4-8 seconds per generation, Omni 4-10 seconds.
- Upscaling: Omni upscales to 1080p/4K inside the generation task; Veo uses separate upscale endpoints at $0.05 (1080p) and $0.50 (4K).
What each model is
Gemini Omni 1.1 Flash is Google's multimodal video model: text, images, video, and audio all sit in the model context. Its defining features are reference-driven generation (name a character, reuse it), an edit endpoint that rewrites an uploaded video from one instruction, and a 10-second context window that keeps scenes consistent across follow-up edits. It replaced the May 2026 Gemini Omni Flash; the full review is in our launch coverage.
Veo 3.1 is Google's established text-to-video and image-to-video model, live since November 2025. It ships in three quality tiers (Lite, Fast, Quality) plus a lower-priority Lite Relaxed, with the most complete production toolkit in the category: start and end frames on every tier, an Extend endpoint that continues a finished video, and dedicated 1080p/4K upscalers. Full pricing breakdown in Veo 3.1 API pricing.
Specs side by side
| Gemini Omni 1.1 Flash | Veo 3.1 | |
|---|---|---|
| Duration per generation | 4, 6, 8, or 10s | 4, 6, or 8s |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16 |
| Audio | Native, same file | Native, same file |
| Resolution | 360p/720p native; 1080p/4K in-task upscale | 720p; separate 1080p/4K upscale endpoints |
| Draft tier | 360p at half price | Lite Relaxed at $0.075 |
| Reference images | Up to 7 | Up to 3 (Fast, Lite, Lite Relaxed only) |
| Character references | Up to 3 named characters, 10 images each | No |
| Start frame | Yes | Yes (every tier) |
| End frame | Yes | Yes (every tier) |
| Extend a finished video | No | Yes, Extend endpoint |
| Edit an uploaded video | Yes, dedicated endpoint (≤30s, ≤1GB source) | No |
| Voice presets | Yes (generation and per-character in edit) | Yes |
| Seed | Yes | No documented seed |
The table is the argument. Omni is built around references and editing, and with keyframes now on its endpoints it covers most of Veo's control surface too; what it cannot do is continue a finished video. Veo's remaining exclusives are Extend and the cheapest draft tier.
Gemini Omni 1.1 Flash vs Veo 3.1 pricing
Unifically bills both models per generated video, pay-per-use, no subscription. All prices verified August 28, 2026 against the live pricing endpoint.
| Route | Price per video |
|---|---|
| Veo 3.1 Lite Relaxed | $0.075 |
| Veo 3.1 Lite | $0.15 |
| Veo 3.1 Fast | $0.30 |
| Veo 3.1 Quality | $0.60 |
| Veo 3.1 Upscale 1080p / 4K | $0.05 / $0.50 |
| Gemini Omni 1.1 Flash generate, 4s / 6s / 8s / 10s | $0.45 / $0.60 / $0.75 / $0.90 |
| Gemini Omni 1.1 Flash 360p draft, 4s-10s | $0.225-$0.45 |
| Gemini Omni 1.1 Flash edit | $1.20 |
Veo is cheaper at every comparable tier, including drafts: a Lite Relaxed video is $0.075 against Omni's cheapest 360p draft at $0.225. Omni claws it back at the top end. Its price is duration-keyed and resolution-free from 720p up, so a 10-second 4K video is $0.90 in one task; the same deliverable on Veo Quality is $1.10 across two tasks ($0.60 generate plus $0.50 upscale) and stops at 8 seconds. Both models now support a cheap-draft loop: iterate on Veo Lite Relaxed or Omni 360p until the prompt is right, then rerun the winner at final quality (Omni's 360p output can only be upscaled to 720p, so treat it strictly as an iteration tool).
Arena benchmarks: Gemini Omni 1.1 Flash vs Veo 3.1
Arena leaderboard positions as of August 28, 2026:
| Board | Gemini Omni 1.1 Flash | Best Veo 3.1 row |
|---|---|---|
| Text-to-video | #1, 1515 Elo (1,762 votes) | #10, 1364 Elo (veo-3.1 Quality tier, 13,705 votes) |
| Image-to-video | #2, 1488 Elo (3,720 votes) | #11, 1398 Elo (25,114 votes) |

Context for those numbers, so the table does not oversell:
- Omni 1.1's text-to-video rating is one day old on 1,762 votes with a rank confidence interval of 1-3. It will move. Veo's ratings sit on tens of thousands of votes and are stable.
- Veo 3.1 is not #1 anywhere on the current video boards, and that is the honest state of things: SeeDance, MiniMax, and the Omni family have passed it on blind preference since spring. It remains ahead of most of the field on production control, which the Arena boards do not measure.
- On image-to-video, Omni 1.1 is second to MiniMax Hailuo H3 (1494), not first. If pure image-to-video preference is your only criterion, H3 deserves a look over both Google models.
Which one should you use?
Our call, having run both: default to Omni 1.1 Flash for new work, keep Veo 3.1 for structured pipelines.
Pick Gemini Omni 1.1 Flash when:
- Output quality per prompt is the priority; the Elo gap is large enough to see in blind side-by-sides.
- Your workflow is reference-heavy: consistent characters, product stills, style references.
- You need to edit existing videos with a prompt. Nothing else on the platform does this.
- You need keyframe transitions at top visual quality: both models take start and end frames, and Omni ranks far higher on output.
- You want 4K without managing a second upscale task.
Pick Veo 3.1 when:
- You chain videos into longer narratives with Extend. Omni stops at 10 seconds per generation.
- You iterate heavily before committing. $0.075 drafts are a third of Omni's cheapest draft, and at volume that compounds.
- You want a known 24 FPS deliverable spec that has been stable in production for months.
The wrong reason to pick Veo is loyalty to the familiar option: for straight text-to-video, reference-to-video, and now keyframe work, the blind-preference gap is real and the request bodies are similar enough that switching costs an afternoon. The wrong reason to pick Omni is the leaderboard alone: if your pipeline depends on Extend or draft volume, a higher Elo does not compensate.
How to access the Gemini Omni and Veo 3.1 APIs
Both run through the same async task API. Omni 1.1:
curl -X POST https://api.unifically.com/v1/tasks \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "google/gemini-omni-flash-1.1-video",
"input": {
"prompt": "A slow dolly across a rain-soaked neon street, reflections rippling",
"duration": 8,
"aspect_ratio": "16:9",
"resolution": "1080p"
}
}'
Veo 3.1 Fast with a start and end frame:
curl -X POST https://api.unifically.com/v1/tasks \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "google/veo-3.1-fast",
"input": {
"prompt": "The camera pushes in as the scene transitions from day to night",
"start_frame_url": "START_IMAGE_URL",
"end_frame_url": "END_IMAGE_URL",
"aspect_ratio": "16:9"
}
}'
Both return a task_id to poll at GET /v1/tasks/<task_id>. Try either in the browser first: Omni 1.1 playground, Veo 3.1 playground.
Frequently asked questions
Is Gemini Omni 1.1 Flash better than Veo 3.1?
On blind human preference, yes: as of August 28, 2026 it leads Veo 3.1's best Arena row by about 150 Elo on text-to-video (1515 vs 1364) and 90 on image-to-video (1488 vs 1398). Veo 3.1 still wins on production control: end frames, the Extend endpoint, and cheaper drafts.
Is Veo 3.1 cheaper than Gemini Omni 1.1 Flash?
Yes, at every tier. Veo 3.1 runs $0.075 (Lite Relaxed) to $0.60 (Quality) per video, while Omni 1.1 generation runs $0.45 (4s) to $0.90 (10s). Omni narrows the gap on 4K output, which costs nothing extra there but adds $0.50 on Veo. Check the pricing page for current rates on both.
Can Veo 3.1 edit an existing video?
No. Veo 3.1 generates from text, frames, and reference images, and can extend its own finished videos, but it cannot take an uploaded video and modify it. Video editing is Gemini Omni 1.1 Flash's territory via google/gemini-omni-flash-1.1-video-edit.
Does Gemini Omni 1.1 Flash support first and last frames like Veo 3.1?
Yes. Both models accept a start and an end frame: Omni 1.1 via start_image_url and end_image_url on its generate endpoint, Veo 3.1 on every tier. For keyframe transitions the choice now comes down to output quality and price rather than capability, and on quality Omni holds the large Arena lead.
Which is better for image-to-video?
Neither is the outright leader. On the Arena image-to-video board (August 28, 2026), MiniMax Hailuo H3 ranks #1 at 1494 Elo, Gemini Omni 1.1 Flash #2 at 1488, and Veo 3.1's best row #11 at 1398. Between the two Google models, Omni is clearly ahead; across the market, test H3 too.
Are both models available through one API?
Yes. Unifically serves both through the same /v1/tasks endpoint and API key, so you can route per job: Omni for reference-led generations and edits, Veo for keyframe transitions and extended narratives.
Related reading
- Gemini Omni 1.1 Flash API: pricing and how to access it: full launch review.
- Veo 3.1 API pricing: every Veo tier, endpoint, and rate.
- Gemini Omni Flash 1.1 model page and Veo 3.1 model page: playgrounds and parameters.
- Veo 3.1 vs SeeDance 2.0: how Veo stacks up outside Google's lineup.
Arena numbers move fast in a launch window; we re-extract them on every edit to this post. Keyframes and 360p drafts already reached Omni's endpoints in launch week and are reflected above; the next thing that would move the verdict is scene extension, which would erode Veo's Extend advantage. We will update when it lands.



