7 papers
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
Quanjian Song, Xinyu Wang, Donghao Zhou +3
Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. While recent advances have improved…
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
Xin Wu, Zhixuan Liang, Yue Ma +3
Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable cha…
Visual Language Hypothesis
Xiu Li
We study visual representation learning from a structural and topological perspective. We begin from a single hypothesis: that visual understanding presupposes a semantic language…
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
Xinyu Shao, Yanzhe Tang, Pengwei Xie +6
Many language-guided robotic systems rely on collapsing spatial reasoning into discrete points, making them brittle to perceptual noise and semantic ambiguity. To address this chal…
TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
Kechen Jiao, Zhirui Fang, Jiahao Liu +9
Using effective generalization capabilities of vision language models (VLMs) in context-specific dynamic tasks for embodied artificial intelligence remains a significant challenge.…
Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning
Chang Tian, Matthew B. Blaschko, Mingzhe Xing +3
Reinforcement learning (RL) has become a key technique for enhancing the reasoning abilities of large language models (LLMs), with policy-gradient algorithms dominating the post-tr…