Zhitong Finance App News, Yun Zhisheng (09678) issued an announcement. The company officially released the U2 Flash, a next-generation high-density intelligent model based on enhanced training based on the U2 general base model. Compared with the previous U2 model, U2-Flash improved significantly in terms of task completion quality and overall efficiency: in the DeepSWE v1.1 list representing programming ability, the U2-Flash score doubled compared to the previous model, surpassed models such as GLM5.3-Flash and DeepSEEK-v4-Pro-0813 with a score of 64.6; the TerminalBench 3.0 score was greatly increased to 24.3 points, surpassing the K3 trillion parameter level model; SW-bench Pro's score reached 61.6 points , an increase of 10.5 points over the previous generation. At the same time, the model achieves comprehensive optimization in terms of end-to-end execution efficiency and inference usage costs, reducing the number of agent task iteration steps by 20% to 30%, shortening the task execution cycle by 35%, and reducing token consumption by 20% to 30%.
U2-Flash uses a sparse hybrid expert (MoE) architecture. The total parameters are about 266B, the single inference activation is about 10B parameters, the average initial response time is controlled within 3 seconds, and the peak output throughput is up to 300 tokens/s. The model unifies competencies in various fields such as programming, intelligence, mathematical reasoning, and instruction following into the same set of weights to achieve the quality of completion of main tasks at a scale far smaller than that of similar dense models. The model has implicit thinking and continuous state reasoning mechanisms, and encapsulates the inference intensity as a four-level controllable interface, and outputs a complete and readable inference chain at high intensity levels to ensure that the inference process is verifiable and traceable. In addition, U2-Flash has completed systematic adaptation to mainstream domestic computing power platforms, providing more flexible computing power options for large-scale deployment.
In terms of training methods, U2-Flash has established an autonomous closed-loop mechanism where the model is deeply involved in its own training, which is the company's first step in the direction of recursive self-improvement (RSI): the model can participate in task generation, trajectory analysis and error correction, and independently construct a high-quality SWE task set with a scale of nearly 100,000; through innovations such as asynchronous agent RL, online strategy distillation of multiple self-training teacher models, and adaptive dynamic task generation, the number of effective training trajectories is increased by about 60%, and the number of training steps can be reduced by about 55%. All independent adjustments are made within a manually set sandbox environment and verification standards, and complete, rollback records are kept.
The release of U2-Flash marks a new stage in the company's big model technology from “human-led model training” to “models participate in their own evolution.” Through a recursive self-improvement mechanism, the model can continuously self-verify, self-correct, and self-enhance in real tasks, laying the foundation for long-term compound interest improvement of model capabilities. In the future, the company will adhere to the technical route of prioritizing verification, full traceability, and clear boundaries, steadily expand the scope for models to participate in their own ability growth on the premise of ensuring safety and control, promote large-scale implementation of high-density intelligence in more real production scenarios, and create long-term value for shareholders.