9 papers
Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy
Haoran Ding, Liang Ma, Yaxun Yang +7
Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as back…
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
Yu Sun, Meng Cao, Yang Ping +24
Vision-Language-Action (VLA) models and world-action models have emerged as central paradigms for general-purpose robotic intelligence, yet their empirical progress remains constra…
World2Act: Latent Action Post-Training from World Model Dynamics
An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6
World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
Yangming Shi, Shixiang Zhu, Tao Shen +14
We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the m…
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
Mingfei Han, Haihong Hao, Liang Ma +6
Vision-and-Language Navigation (VLN) has long been constrained by the limited diversity and scalability of simulator-curated datasets, which fail to capture the complexity of real-…
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
Liang Ma, Jiajun Wen, Min Lin +12
While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particul…