10 papers
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
Jiawen Wen, Penglei Sun, Wenjie Zhang +4
As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. Howev…
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
Taowen Wang, Zikang Xie, Bin Yang +13
Humanoid robots promise whole-body interaction in human-centered environments, but scalable policy learning remains difficult because task-level decision-making and whole-body dyna…
Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis
Weisheng Xu, Jian Li, Yi Gu +12
Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-bas…
Beyond Viewpoint Generalization: What Multi-View Demonstrations Offer and How to Synthesize Them for Robot Manipulation?
Boyang Cai, Qiwei Liang, Jiawei Li +11
Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling be…
DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
Jiaqi Xiong, Yunjia Qi, Qi Cao +6
Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process a…
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
Jiaxi Zhang, Yunheng Wang, Wei Lu +8
3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamental to embodied AI applications. Although foundation models enable open-…