works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

Motubrain: An Advanced World Action Model for Robot Control

Motubrain Team, Chendong Xiang, Fan Bao +17

Motubrain is a unified world action model that jointly learns video and robot actions using a UniDiffuser and Mixture-of-Transformers architecture, enabling policy learning, world…

cs.CV2026

GEM: Generative Supervision Helps Embodied Intelligence

Ruowen Zhao, Bangguo Li, Zuyan Liu +9

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a si…

cs.CV2026

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Tencent Robotics X, HY Vision Team, : +20

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) an…

cs.CV2025

Motus: A Unified Latent Action World Model

Hongzhe Bi, Hengkai Tan, Shenghao Xie +13

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation pr…

cs.CV2025

NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks

Junliang Ye, Shenghao Xie, Ruowen Zhao +5

3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient, inconsistent, and often fail to prese…

cs.CV2025

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Junliang Ye, Zhengyi Wang, Ruowen Zhao +2

Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities r…