Orply.

MiniMax H3 Unifies Video, Image, and Audio References in One Generation

Alec WilcockElevenLabsThursday, September 10, 20264 min read

MiniMax argues that video generation’s separate tools for motion, subjects, frames and audio can be consolidated into H3, a model that takes text, image, video and audio references in a single generation. Presented by ElevenLabs in its ElevenCreative platform, H3 produces five- to 15-second clips with generated audio, timecoded multi-shot prompts and a 2K regeneration pass that retains the original instructions and references. H3 Max, built by fal from MiniMax’s published weights, trades that 2K option for a claimed roughly three-second generation time for a five-second 768p clip.

H3 treats multimodal reference as one generation problem

MiniMax H3 is positioned as a general-purpose video model rather than a collection of specialized tools. It accepts text, images, video, and audio in a single generation, producing a video with accompanying audio for clips up to 15 seconds. Inside ElevenCreative, the standard H3 option is presented as supporting 2K output, video and audio references, and audio-backed generation.

MiniMax’s premise is that video creation has become fragmented: one model for text-to-video, another for start and end frames, another for subject consistency, another for motion. H3 is intended to consolidate those jobs. That is not an entirely new category—Seedance and Gemini Omni Flash have begun addressing the same problem—but H3 arrives as another strong competitor in it.

The practical interface is deliberately simple. A creator can attach assets and assign each one a role in ordinary prompt language: take camera movement from a video, use an image as the character reference, or match a voice to an audio clip. The clearer the job assigned to each reference, the more predictable the generation is meant to be. References should not be added as undifferentiated context.

The available reference budget is substantial but bounded:

Reference typeMaximum attachmentsAdditional constraint
Images912 total files across all reference types
Video clips3Each clip must run from 2 to 15 seconds
Audio clips3Each clip must run from 2 to 15 seconds
MiniMax H3’s reference limits within a single generation

Start and end frames are supported, but they are a separate mode: they cannot be combined with image, video, or audio references. That creates a specific consistency constraint. If an object does not appear in either boundary frame, it cannot be supplied separately as a reference for that generation.

Timecoded prompts turn a short clip into a sequence of scenes

H3 generates audio alongside the video, with MiniMax claiming stable dialogue in 11 languages. Clips run from five to 15 seconds at 24 frames per second.

Its more consequential creative feature is native multi-shot generation. Rather than generating isolated clips and stitching them together afterward, users can write cuts directly into the prompt with time ranges. A 15-second prompt might specify one shot from zero to five seconds, a second from five to 10, and a final shot from 10 to 15.

The illustrated prompt stages a tailor’s atelier as a three-shot sequence: a 100mm macro shot of white chalk marking dove-grey wool; a 65mm medium shot of a tailor pinning a jacket on a client; then a wide 35mm shot of the client turning before a mirror. It specifies the lighting, muted palette, fine grain, and diegetic audio—chalk, pins, fabric, and hush—with no music.

That structure asks the model to retain objects, characters, and locations across changes in framing and action. H3’s intended use is not merely a single uninterrupted scene, but a compact visual sequence whose cuts are authored in text while continuity remains within one generation.

The 2K option regenerates the clip with its original context

H3 generates initially at 480p or 768p. Its advertised 2K output begins with the 768p result, then makes a higher-resolution pass with the original prompt and reference assets still in context.

A MiniMax technical-report excerpt calls this method “H3-in-context Regeneration.” Instead of applying a conventional dedicated super-resolution module, H3’s base model regenerates its own low-resolution output while again receiving the multimodal instructions that guided the first generation.

2K
H3’s highest advertised output resolution, produced from a 768p generation with prompt and references retained in context

MiniMax says this approach can recover details such as small text and fine visual elements that a conventional super-resolution system would have to infer from pixels alone. The significance is not simply that the clip is sharpened after generation. The higher-resolution output is presented as a second generative step informed by the same text, images, video, and audio references that established the intended result.

H3 Max exchanges 2K output for rapid iteration

H3 Max is not described as a larger version of H3. It is a faster implementation, with output topping out at 768p.

Its existence is tied to MiniMax publishing H3’s model weights—the model files that others can download and build on. According to the presentation, fal used those weights to rebuild H3 for speed and produced H3 Max. The performance claim is a five-second video in roughly three seconds of generation time.

~3 seconds
Claimed generation time for a five-second H3 Max video

The trade-off is explicit: H3 Max does not offer standard H3’s 2K regeneration path. For workflows where iteration speed matters more than final resolution, that limitation is the point of the product. A creator can test prompts and ideas rapidly at up to 768p, then compare results with H3 when the higher-resolution pass is warranted.

Both models are available in ElevenCreative’s Image & Video model picker. Standard H3 is the option for multimodal references and 2K output; H3 Max is the speed-first alternative when 768p is sufficient.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free