The Zhitong Finance App learned that Huachuang Securities released a research report saying that recently, MiniMax officially opened the H3 basic model weight. The bank indicates that the key to H3 is not only improving resolution, but also transforming complex full-modal generation into a low-cost product that can be used on a large scale through model, data, and infrastructure collaboration. Competition for video models is shifting from single-shot material generation to understanding complete creative intent and participating in production processes such as advertising, e-commerce, gaming, and design. The domestic big model has entered a new stage where reasoning efficiency and open source ecology are equally important. Industry competition has moved from simply improving intelligence to a comprehensive competition of model layering, full modal generation, agent execution ability, and unit task cost.
The main views of Huachuang Securities are as follows:
Tasks and modes are unified, and the video model is moving from a generation tool to a general creation system
H3 can unify the multi-modal context of text, images, video, and audio, and generate video up to 15 seconds long, up to 2K resolution, with native stereo audio. Unlike traditional models, which split Wensheng video, motion reference, sound reference, and video editing into multiple independent tasks, H3 allows users to directly describe relationships between materials in natural language, and complete subject maintenance, motion migration, sound references, and content editing within the same model. The bank believes that competition for video models is shifting from generating single footage to understanding complete creative intent and participating in production processes such as advertising, e-commerce, games, and design.
The architecture is comprehensive and the service tasks are generalized, and efficiency advantages form an important foundation for H3 commercialization
H3 establishes relationships between different materials and target content through contextual Omni Representation, and uses a 33B parameter single-stream H3-Omni Transformer to jointly predict video and audio; H3-VAE's high compression ratio brings 4x sequence length benefits, and understanding and generating heterogeneous training architectures increases end-to-end training throughput by nearly 30%. In the 2K output process, H3 re-enters the 768p results with the original context back into the base model to improve the ability to restore small text and detailed content. The bank believes that the key to H3 is not only improving resolution, but also transforming complex full-mode generation into a low-cost product that can be used on a large scale through model, data, and infra collaboration.
The open foundation model accelerates ecological diffusion, and MiniMax is exploring open source and commercialization collaboration
With H3-Base as the core, this release supports Wensheng audio and video, first-and-end frame generation, and full-modal reference generation, and is compatible with mainstream deployment frameworks such as SGLang, VLLM, DiffUsers, and ComfyUI; H3-Context-IR, which is responsible for complex input orchestration, and H3-regenerate-2K, which is responsible for 2K regeneration, are temporarily provided through the official API. Within 24 hours after opening, 9 chip manufacturers including Huawei Ascend, Moore Thread, Mu Xi, Haiguang, AMD, and Intel completed Day 0 adaptation, and more than 100 developer communities, cloud inference platforms, and enterprises completed access. The bank believes that “open basic model+retaining high-value orchestration and high-definition services” can not only expand ecological influence, but also preserve a clear commercial entrance for API calls and enterprise services.
Models have been concentrated and updated in the past two weeks, and performance upgrades and price reductions are accelerating simultaneously
On July 30, OpenAI cut GPT-5.6 Luna prices by 80% and Terra by 20%. After adjustments, Luna is $0.20 per million tokens input and $1.20 per million tokens output; Terra's API price is $2 per million tokens and $12 per million tokens; the price of Sol remains unchanged. Anthropic launched Claude Opus 5 on July 24, which supports 1M context and 128K output, and is priced at $5/25, focusing on strengthening complex agentic coding and enterprise tasks. Domestically, Qwen launched 2.4T parameters and Qwen3.8-Max with 1M context on August 3; DeepSEEK-v4-Flash-0731 was launched on August 1, focusing on high concurrency and low cost scenarios with 284B total parameters, 13B activation parameters, and 1M context. The bank believes that industry competition has moved from simply improving the upper boundary of intelligence to a comprehensive competition of model layering, full modal generation, agent execution ability, and unit task cost.
Risk warning: Technological progress falls short of expectations; model implementation falls short of expectations; commercial implementation falls short of expectations.