Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems, Emerging Technologies & Innovation · Depth: Expert, quick

Summary

Xiaomi-Robotics-U0 is a 38-billion-parameter multimodal autoregressive model designed for unified embodied synthesis, extending foundation image and video generation to complex embodied scenarios. This framework jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation, preserving the generalization of pre-trained world foundation models. It is the first model to support high-quality multi-view scene generation across multiple robot embodiments and introduces structured, controllable embodied transfer for fine-grained editing while maintaining multi-view consistency and interaction dynamics. Xiaomi-Robotics-U0 achieves state-of-the-art results, outperforming GPT-Image-2.0 in human evaluations for embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of pi_0.5 from 36.9% to 63.2% on challenging real-world manipulation tasks. This demonstrates its potential as both an embodied world model and a scalable data engine.

Key takeaway

For Robotics Engineers developing advanced manipulation systems, Xiaomi-Robotics-U0 offers a powerful new approach to data generation and simulation. You should explore this unified embodied synthesis model to create high-quality, multi-view consistent training data, potentially reducing reliance on real-world data collection. Its ability to improve out-of-distribution success rates suggests a significant leap in preparing robots for complex, unpredictable environments.

Key insights

Xiaomi-Robotics-U0 unifies embodied synthesis by extending world foundation models, enabling multi-view consistent robot interaction and generation.

Principles

Method

The model jointly optimizes multiple embodied generation tasks, including text-to-image, scene generation, and embodied transfer, within a single autoregressive framework to adapt world foundation models.

In practice

Topics

Best for: Research Scientist, AI Scientist, Robotics Engineer, Machine Learning Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.