Shen Wan Hongyuan: Humanoid robots enter embodied intelligence, medium- to long-term suggestions for layout along three main lines

Zhitongcaijing · 2d ago

The Zhitong Finance App learned that Shen Wan Hongyuan released a research report saying that humanoid robots have entered a new stage of artificial intelligence. Main body components are the hardware foundation, and mass production drives the release of demand for speed reducers, servos, and sensors; large models are the core constraints of robot intelligence, and algorithms and simulation platforms build software barriers; training data is the core constraint of model iteration. Multiple routes such as real teaching acquisition+first-perspective collection+simulation synthesis data are parallel, and the scarcity of data supply continues to increase. The medium- to long-term proposal is based on three main lines: ① core component manufacturers for humanoid robots; ② physical model and body manufacturers; ③ targets related to robot data acquisition hardware and data services.

Shen Wan Hongyuan's main views are as follows:

The four key elements in the development of the robotics industry are models, data, ontologies, and scenarios

Among them, the model is the robot brain, which is responsible for reasoning decisions; data is the fuel for model training; without high-quality data, the model cannot be effectively trained; the body is the physical carrier of the robot, the physical upper limit of execution ability, and all intelligent instructions are ultimately executed by the body; the scene is the focus of the final commercial implementation, and the source of demand for the robot also determines the robot's task boundaries; the robot must have a scene to generate value. The bank believes that at this stage, models and data are the biggest bottlenecks in the industry. The market has paid less attention and research to this aspect.

model

Mainstream players' models continue to be iterated, and the model architecture has yet to converge. Mainstream players' models continue to be iterated, from PI's 0 to 0.7, Google's RT1 to Gemini Robotics 2, and continuous breakthroughs in cross-ontology generalization, operation fluency, and data generation bottlenecks; the model architecture has not yet converged, but the similarity is high. Most of them are integrated architectures such as VLA, diffusion models, world models, etc., and brain reasoning and cerebellar movement models are layered. Participants showed a trend full of flowers: 1) Big Tech Companies: Google, Byte; 2) Embodied Models & Ontology Companies: PI, Figure, Zhiyuan, Starsea Map, etc.; 3) Big Language Model Vendors: OpenAI, etc.

data

The core fuel for model training is data quantity, quality, and scale-up. Compared to large language models, which have massive Internet data training, embodied intelligence naturally lacks robot trajectory data, so data has become a key bottleneck in model iteration. Current data sources include Internet data, synthetic simulation data, first-person perspective data, remote operation data, etc. Currently, different companies are betting on different data routes and generally use multiple data types, open source+internal data set integration methods. In the model training phase, data producers and acquisition equipment vendors ushered in investment opportunities.

Risk warning: risk of model technology iteration falling short of expectations, risk of bottlenecks in data supply and quality, risk of commercial implementation.