
Google Gemini Omni
Gemini Omni Reference images to video, support 1-5 reference images to generate video.
Gemini Omni Flash 1.1 image-to-video accepts a first frame and an optional last frame to strictly control the start and end of the video, with synchronized native audio and up to 4K resolution.
Explore similar AI models available through the same unified API

Google Gemini Omni
Gemini Omni Reference images to video, support 1-5 reference images to generate video.

Seedance 2.5 (Ultra-low Price)
Seedance 2.5 X Reference to Video combines up to 30 reference images, 10 reference videos and 10 reference audios into a 4-30 second clip at 480p or 720p, priced by the second. Real-person material is supported and passed through as public URLs.
Seedance 2.0 (Ultra-low Price)
Seedance 2.0 X Reference to Video combines up to 9 reference images, 3 reference videos and 3 reference audios into a 4-15 second 720p clip at a flat price per generation. Reference materials are cited in the prompt with @Image1 / @Video1 / @Audio1, and real-person material is supported.

Seedance 2.5
Seedance 2.5 Draft on APIPod is a two-stage workflow. Submit a normal Seedance 2.5 request to get a low-cost 480p draft; once it completes, submit its task ID as drafttaskid to render the 1080p final video. The final video reuses the draft prompt, media, duration, aspect ratio, seed, and audio setting, and runs on the same generation account as the draft.
MiniMax H3 Max
MiniMax H3 Max is the fal post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics. Image to Video animates a required first frame with an optional last frame for first-to-last keyframe control; the output aspect ratio follows the first frame.

Wan3.0 Video Prime
WAN 3.0 Prime Image to Video is the high-speed tier of the WAN 3.0 image-to-video model. It accepts a first frame and an optional last frame to strictly control the start and end of the video, with significantly faster end-to-end generation.

Seedance 2.5(Budget Version)
Seedance 2.5 Reference to Video Model, Full‑Power Version, comes with a material library, no facial distortion, 100% real‑human pass rate, long‑term stability. Priced at approximately 60% of the official rate. Material and Scene Limitations Text Prompts: Chinese text is recommended to be no more than 500 characters; English text is recommended to be no more than 1000 words. Image Input: Supports jpeg, png, webp, bmp, tiff, gif, heic, heif; single‑image size below 30 MB, aspect ratio (0.4, 2.5), width and height range (300px, 6000px). Image Count: No more than 30 images. Video Input: Single‑video duration [2, 30] s; up to 10 reference videos allowed; total duration of all videos shall not exceed 30 s. Audio Input: Single‑audio duration [2, 30] s; up to 10 reference audio clips allowed; total duration of all audio clips shall not exceed 30 s. Audio cannot be input independently; at least one reference image or video shall be provided simultaneously. 🔸Note🔸: Unit price for with‑video input is lower due to different calculation formulas: No video = Unit Price × Output Volume; With video = Unit Price × (Input Volume + Output Volume).