collaborators

18 papers

cs.RO2026

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

Weili Zeng, Yitong Xing, Fulong Liu +10

The paper introduces Enfold, a method that folds the computation of a world-generative model into a predictive representation derived from the current visual scene and language ins…

cs.CV2026

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

Boyu Mi, Mengchen Ma, Yifei Yao +10

The paper introduces REAL, a framework that trains embodied agents for open‑world mobile manipulation using sim‑to‑real consistent environments, hierarchical training, and human‑in…

cs.RO2026

-EqM: Equilibrium Matching for Closed-Loop Vision-Language-Action Control

Huanming Liu, Congsheng Xu, Jianmin Ji +1

Currently, Vision-Language-Action (VLA) models have become the most adopted paradigm for robotic manipulation for its great potential for task generalization. While most generative…

cs.RO2026

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26

Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…

cs.RO2026

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

Zili Lin, Wenyao Zhang, Yuyang Zhang +9

Demonstration augmentation is proposed for cost-efficient data acquisition, but existing methods are fundamentally limited in deformable manipulation due to two challenges: (1) the…

cs.RO2026

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Rongxu Cui, Zongzheng Zhang, Jingrui Pang +11

Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains largely unverified. To address th…