Feedback
AI Ad Video Example
Loading...
minimax h3 video model
Produce 2K video with synchronized audio via the minimax h3 video model — a single engine for text, images, footage, and sound in up to 15 seconds.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.
Key Advantages of the minimax h3 video model for Creators
As MiniMax's open-weight omni-modal model, the minimax h3 video model lives on fal.ai from day one. It ingests text, stills, footage, and audio in one pass, then outputs 2K video with built-in stereo sound for up to 15 seconds. You also get targeted editing, sharp text/UI rendering, and as many as 12 reference inputs per run.
- A Single Unified Input SpaceWith support for up to nine images, three clips, and three audio tracks in one request, the minimax h3 video model blends character, motion, camera, and audio into a seamless output.
- Synced Built-in Stereo SoundEach output from the minimax h3 video model includes original music, spoken dialogue, foley, and ambient effects locked to the edit, plus voice transfer and cloning from audio references.
- Region-Specific Video EditingSwap objects, adjust text, redub lines, or shift daylight to night — the minimax h3 video model changes only the chosen area, leaving everything else unchanged.
Steps to Run the minimax h3 video model
Follow three simple steps to generate 2K video with synced audio through the minimax h3 video model API.
Feature Set of the minimax h3 video model
With three dedicated API endpoints, one unified multimodal context, synced sound, targeted edits, crisp on-screen text, and flexible usage-based pricing, the minimax h3 video model gives you a full 2K production workflow on fal.ai.
Three Dedicated Generation Modes
For the minimax h3 video model, you can use text-to-video, image-to-video with optional start/end frame control, and reference-to-video endpoints, covering all creative workflows.
Up to a Dozen Referenced Media Inputs
Mix nine images, three short clips, and three audio files; the minimax h3 video model extracts character identity, motion, camera style, layout, and pacing from those references.
Clean Text and UI Rendering
Produce crisp titles, end screens, captions, and logo reveals, and animate working interfaces — websites, game menus, HUDs, and kinetic type using the minimax h3 video model.
Long Prompts for Full Scene Control
Add an entire shot list to one request; the minimax h3 video model handles prompts as long as 7,000 characters to direct every part of the scene.
2K Output at 24 Frames Per Second
Generate 2K video with a 1440px shorter edge, up to 15 seconds at 24fps, plus six aspect ratio presets and an adaptive option from the minimax h3 video model.
Usage-Based API Pricing
Access the minimax h3 video model on a serverless, pay-as-you-go basis — no minimum commitments, no monthly plans, and commercial rights to the videos you create.
Frequently Asked Questions About the minimax h3 video model
Answers to frequently asked questions about using the minimax h3 video model on fal.ai.
What is the minimax h3 video model designed to do?
The minimax h3 video model is MiniMax's open-weight, omni-modal generation engine available on fal.ai from launch. It ingests text, images, video, and audio together in a single context, then outputs up to 15 seconds of 2K video with built-in stereo audio.
Which endpoints are available for the minimax h3 video model?
There are three ways to call the minimax h3 video model: text-to-video, image-to-video with optional start/end-frame control, and reference-to-video, which preserves subjects, style, motion, camera choices, and voices from your reference files.
What output sizes and lengths can the minimax h3 video model create?
The minimax h3 video model generates 2K footage (1440px on the short edge) at 24fps, running 5 to 15 seconds. Supported aspect ratios are 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and an adaptive preset.
Can I get sound in the videos produced by the minimax h3 video model?
Absolutely — every video from the minimax h3 video model includes stereo sound: composed music, spoken dialogue, foley, and room tone synced to the cut, plus voice transfer or cloning from reference audio.
What is the maximum number of reference files for a single generation?
You can provide up to twelve files: nine reference pictures, three reference clips of 2–15 seconds each, and three audio tracks of 2–15 seconds each. For the minimax h3 video model, each audio file must be paired with at least one image or video.
Is commercial use allowed for content created with this model?
Yes. Videos created through the fal.ai API using the minimax h3 video model can be used in commercial work, subject to fal.ai's terms of service.
Begin Creating Videos with the minimax h3 video model
Kick off your next project by generating 2K video with synced stereo audio in a single call to the minimax h3 video model — multimodal reference support, precision edits, and flexible API pricing on fal.ai.
