Shanghai's MiniMax open-sourced MiniMax H3 Base, an omni-modal video model that jointly reads text, images, video, and audio and can generate about 4 to 15 second clips with native stereo sound. First/last-frame and reference-conditioned checkpoints are on Hugging Face under a MiniMax community license, while the Context-IR layer and 2K regenerate module stay closed and API-only. Why it matters: an open Chinese video stack is now competing near the top of public generation and editing leaderboards. Caveat: full 2K still needs the closed module, and the license restricts commercial use for larger firms.