Chinese AI firm MiniMax has released H3, a video-generation model that produces clips of up to 15 seconds in 2K resolution with native stereo sound, according to Reuters reporting by Eduardo Baptista, surfaced via Techmeme.

Reuters reports the model, released Friday, can process text, images and other inputs, and that MiniMax plans to release H3's weights within days — meaning outside developers would be able to download and run the model themselves rather than only accessing it through a company-controlled service.

MarkTechPost frames H3 as something broader than a video generator with extras bolted on. In its description, MiniMax calls H3 a general-purpose multimodal generation model that reads text, images, video and audio as a single unified context, and returns video with native stereo audio. The distinction matters: in many systems, sound is generated separately and stitched onto finished footage. Treating audio as part of the same generation process is a different architectural bet.

MiniMax trades publicly in Hong Kong under the ticker 0100.HK, per the Reuters item.

Two details carry the most weight here. The first is resolution and audio quality together — 2K with stereo sound pushes AI-generated clips closer to the threshold where they can slot into real production work without obvious tells. The second is the weights release. Open weights mean researchers, startups and hobbyists can inspect, fine-tune and deploy the model independently, which tends to accelerate both useful applications and harder-to-police uses.

This matters because it puts a high-fidelity, audio-native video generator into the hands of anyone who can download it, and it lands another marker in the fast-moving competition between Chinese and American AI labs over who sets the pace in generative video.