Recently, the Stardust Smart Pedestal Model Team released SmoothRL, an online reinforcement learning framework that can be executed asynchronously. The solution is that after asynchronous inference of large models becomes the norm in actual deployment, online RL adapted to it should learn from what actions. SmoothRL verified this type of asynchronous online RL for the first time in a real high-dynamic throwing mission, further proving that online learning can not only be used for high-precision correction, but can also adapt to continuous acceleration, accurate release, and unstoppable dynamic operation.

Zhitongcaijing · 2d ago
Recently, the Stardust Smart Pedestal Model Team released SmoothRL, an online reinforcement learning framework that can be executed asynchronously. The solution is that after asynchronous inference of large models becomes the norm in actual deployment, online RL adapted to it should learn from what actions. SmoothRL verified this type of asynchronous online RL for the first time in a real high-dynamic throwing mission, further proving that online learning can not only be used for high-precision correction, but can also adapt to continuous acceleration, accurate release, and unstoppable dynamic operation.