Luo Fuli, head of the Xiaomi Mimo team, published an article on the X platform to reveal for the first time the progress of reinforcement learning training for Xiaomi's new model MiMO-v2.6. Luo Fuli said that in the past six months, the team has been conducting research on the question of how far reinforcement learning can be extended. Currently, MiMO-v2.6 is in the process of reinforcement learning, and the team is expanding the scale in three directions, including computational resources, training environment and task system, and reward evaluation mechanisms.

Zhitongcaijing · 2d ago
Luo Fuli, head of the Xiaomi Mimo team, published an article on the X platform to reveal for the first time the progress of reinforcement learning training for Xiaomi's new model MiMO-v2.6. Luo Fuli said that in the past six months, the team has been conducting research on the question of how far reinforcement learning can be extended. Currently, MiMO-v2.6 is in the process of reinforcement learning, and the team is expanding the scale in three directions, including computational resources, training environment and task system, and reward evaluation mechanisms.