Ant Reimbo LingBot-World 2.0 re-open source three models

Zhitongcaijing · 2d ago

The Zhitong Finance App learned that after opening the LingBotWorld 2.0 14B model in July of this year, today (September 11), Ant Lingbo Technology will further open three models: LingBot-World 2.0 Small (1.3B), LingBot-World 2.0 Bidirectional, and LingBOT-World 2.0 Causal Pretrain.

This time, Ant Lingbo Technology first wants to solve a more direct problem: when the world model can be continuously generated over a long period of time, respond to operations in real time, and continuously create new events, how can more developers actually run it?

The 1.3B small model targets consumer-grade single-card GPUs and can achieve real-time world generation on a single card; Bidirectional and Causal Pretrain further open up model distillation and causal pre-training capabilities. Ant Lingbo Technology said it hopes that LingBot-World 2.0 can not only be experienced, but also easier to operate, research, and improve.

LingBot-World 2.0 is a real-time interactive world model. Previously, in the 14B main model released in July, emphasis was placed on promoting long-term timing stability, real-time generation, and interaction capabilities. The model supports continuous hourly generation, and can respond to character actions such as attacks, archery, casting, and shooting, as well as environmental events such as snow and rain. Agentic Interaction Harness, composed of Pilot Agent and Director Agent, collaborates to generate character behavior and environmental content, so that the world continues to evolve as it explores. With higher computing power and corresponding inference configurations, LingBotWorld 2.0 can further achieve a 720P/60FPS high-definition real-time experience.

The open source LingBot-World 2.0 Small uses a 1.3B parameter scale, is designed for consumer-grade single-card GPUs, and can be generated in real time on a single card.

In addition, Ant Lingbo simultaneously opens the LingBot-World 2.0 Bidirectional, or bidirectional teacher model. It is aimed at real-time world model training and provides a high-quality distillation source. High-quality generation usually requires multi-step denoising, and computational costs are high, making it difficult to directly meet real-time interaction requirements. Through distillation, the teacher's multi-step generation trajectory can be compressed to a less-step student, and while reducing generation costs, the dynamic ability of motion conditions learned during the pre-training phase can be preserved as much as possible.

The third model released this time is the LingBot-World 2.0 Causal Pretrain. It provides a foundation for causal pre-training for post-training, evaluation, and scenario adaptation. Causal Pretrain supports two types of input: camera pose and text prompt (prompt), allowing users to control movement and perspective, as well as change local events and world states through text.