Draft Speed
60% faster
Official 360p preview claim vs 720p — at roughly ⅓ the cost
Gemini Omni Flash 1.1 is Google's lightweight preview model for fast, high-volume video generation and editing. It turns text, images and video into clips with synchronized native audio, supports keyframe interpolation and scene extension, and renders from quick 360p drafts up to 4K finals.
Gemini Omni Flash 1.1 is Google's lightweight preview model for fast, high-volume video generation. Text-to-video turns a prompt into a short clip with synchronized native audio at up to 4K resolution, at roughly a third of the cost of the standard tier.
Gemini Omni Flash 1.1 image-to-video accepts a first frame and an optional last frame to strictly control the start and end of the video, with synchronized native audio and up to 4K resolution.
Gemini Omni Flash 1.1 reference-to-video accepts 1-7 reference images to keep characters and styles consistent across shots, with synchronized native audio and up to 4K resolution.
Gemini Omni Flash 1.1 video extend continues an existing clip from its first 10 seconds, optionally guided by up to 5 reference images and a prompt.
Google · Gemini Omni Flash 1.1
Google's lightweight video model for high-volume generation — iterate at 360p speed, finish at 4K.
Gemini Omni Flash 1.1 turns text, images and video into clips with synchronized native audio. Draft fast at roughly a third of the cost, pin first and last frames for exact camera moves, and extend scenes with a 10-second context window. One unified API, four task modes.
Draft Speed
60% faster
Official 360p preview claim vs 720p — at roughly ⅓ the cost
Scene Extension
10s context
Continuations read 10 seconds of prior footage for consistent cuts
Task Modes
4
Text-to-video, first/last frame, reference and extension tasks
APIPod Price
From $0.45
Per finished video, any 4-10s duration at 360p
Demo clips from the Gemini Omni model family. Shared showcase media with the standard series — same model lineage, same native-audio pipeline.
Both tiers run on the same APIPod unified API. Flash 1.1 is the newer, faster model generation built for iteration volume; the standard Omni tier stays the budget pick for flexible, low-cost clips.
Same family, same unified API — different trade-offs between speed, control and price.
| Specification | Omni Flash 1.1 | Omni Standard |
|---|---|---|
| Model generation | 1.1 (newest, GA Aug 2026) | 1.0 |
| Duration | 4 / 6 / 8 / 10 seconds | 1-30 seconds, any value |
| Resolution | 360p / 720p / 1080p / 4K (default 720p) | 720p / 1080p |
| Aspect ratio | 16:9, 9:16 | 16:9, 9:16, 1:1 |
| Reference images | Up to 7 | Up to 5 |
| Video extension | Continues from the first 10s of source footage | Continues from the final frame |
| Audio | Native synchronized audio | Native synchronized audio |
| APIPod price | $0.45 - $4.50 per video by resolution | $0.25 flat per video |
Both tiers are served through the same /v1/videos/generations endpoint with task polling — switching models is a one-line change.
Transparent Pricing
One price per finished video. Duration is free within the 4-10 second range — a 4-second draft and a 10-second cut cost the same at the same resolution.
| Resolution | Price per video |
|---|---|
| 360p — draft tier | $0.45 |
| 720p — default | $1.35 |
| 1080p | $2.25 |
| 4K | $4.50 |
Reference workflow: storyboard the whole sequence in 360p at $0.45 a shot, then regenerate only the keepers at 4K.
For comparison, official Gemini API list pricing is $0.03 / $0.10 / $0.15 / $0.30 per second at 360p / 720p / 1080p / 4K. APIPod pricing is anchored to a 10-second worst case, so short clips are never penalized.
Every mode is an async task on POST /v1/videos/generations — submit, poll, retrieve the video URL.
gemini-omni-flash-1.1-t2v
Prompt-only generation with synchronized native audio. Pick 4-10 seconds and any resolution from 360p to 4K.
Input: prompt
gemini-omni-flash-1.1-i2v
Keyframe interpolation: supply a first frame and an optional last frame, and the model generates a smooth transition between them — loops, orbits and dolly moves.
Input: prompt + 1-2 frame images
gemini-omni-flash-1.1-r2v
Up to 7 reference images keep characters, outfits and art direction consistent across every shot of a sequence.
Input: prompt + 1-7 reference images
gemini-omni-flash-1.1-extend
Continue an existing clip from its first 10 seconds — the model reads the full context window for consistent motion and seams, optionally guided by reference images.
Input: prompt + source video + optional references
Prompt Library
Real prompt patterns that play to Flash 1.1's strengths — native audio, physics, keyframe control and character consistency. Copy them into the API or console playground.
A single prompt carries scene, camera motion and sound design — the model renders the shot with synchronized ambient audio.
Prompt
A cinematic tracking shot through a rain-soaked neon street at night, reflections shimmering on wet asphalt, the camera gliding past storefronts, ambient rain and distant traffic sound.
Pin a sunrise photo as the first frame and a starry-night render as the last — Flash 1.1 interpolates a continuous day-to-night transition.
Prompt
Smooth transition from the first frame to the last frame: the sunrise over the mountain ridge gradually becomes a starry night, clouds drifting, the camera slowly orbiting to the right.
The same reference character stays consistent across three different environments — outfit, stride and lighting adapt to each city.
Prompt
Using the reference images, the same character walks through three cities — Tokyo, Paris and New York — keeping her outfit and stride consistent, documentary style with ambient street audio.
Feed the generated establishing shot back in as source footage: the 10-second context window keeps motion and lighting continuous across the cut.
Prompt
Continue the scene: the explorer steps through the doorway into the hidden temple, torchlight flickering across the stone walls, suspenseful score swelling.
Demo media is shared with the Gemini Omni standard series (same model family, official showcase clips).
Quickstart
One async endpoint for all four modes. Submit the task, poll the status, pick up the video URL — no SDK required.
Full parameter reference (duration / resolution / reference image limits) lives in the docs for each mode.
import os, requests
payload = {
"model": "gemini-omni-flash-1.1-t2v", # or -i2v / -r2v / -extend
"prompt": "A cinematic tracking shot through a rain-soaked neon street...",
"duration": 8, # 4 / 6 / 8 / 10
"resolution": "1080p", # 360p / 720p / 1080p / 4k
"aspect_ratio": "16:9" # or 9:16
}
resp = requests.post(
"https://api.apipod.ai/v1/videos/generations",
json=payload,
headers={"Authorization": f"Bearer {os.environ['APIPOD_API_KEY']}"},
)
task_id = resp.json()["data"]["task_id"]
# poll GET /v1/videos/status/{task_id} until status == "completed"Flash 1.1 is Google's lightweight preview-tier model for fast, high-volume video generation and editing, generally available via the Gemini API since August 2026. It generates clips with synchronized native audio, supports keyframe interpolation between a first and last frame, and renders from 360p up to 4K. On APIPod it is available through the same unified async video API as every other model.
Flash 1.1 is the newer model generation: it adds a 10-second context for scene extension, 4K output, up to 7 reference images and cheap 360p drafts. The standard Omni tier keeps the lowest flat price ($0.25 per video) and flexible 1-30 second durations. A common pattern is to draft and storyboard on Flash 360p, and use Omni standard when you just need inexpensive straightforward clips.
Flash 1.1 supports 4, 6, 8 or 10 second clips at 360p, 720p (default), 1080p or 4K, in 16:9 or 9:16. Pricing is per video by resolution — any duration within 4-10 seconds costs the same at a given resolution.
Use the gemini-omni-flash-1.1-extend model with a source video. The model reads the first 10 seconds of the source as context — instead of just the final frame — and generates a continuation that keeps motion, lighting and audio consistent. You can attach up to 5 reference images and a prompt to steer the direction of the continuation.
Yes. Flash 1.1 generates synchronized native audio — dialogue, ambience and score — alongside the video. Audio is experimental per Google's notes and may occasionally be unavailable on individual outputs.
Per finished video, by resolution: $0.45 at 360p, $1.35 at 720p, $2.25 at 1080p and $4.50 at 4K. Duration is free within the 4-10 second range. Failed generations are not billed.
POST https://api.apipod.ai/v1/videos/generations with the model set to gemini-omni-flash-1.1-t2v, -i2v, -r2v or -extend. The endpoint returns a task_id; poll GET /v1/videos/status/{task_id} until the status is completed and read the video URL from the result. Full request/response reference is in the APIPod docs.
Draft fast, finish at 4K, and keep native audio in every cut — on the same unified API as the rest of your stack.