From the 1 of 15 linked papers with an AI index.
15 papers
Semantic Anchoring for Robotic Action Representations
Yuan Xu, Youheng Shi, Chengyang Li +2
The paper studies how fine‑tuning vision‑language‑action models for robots can degrade the semantic structure of their action representations, and proposes a plug‑and‑play anchorin…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han, Jiangran Lyu +13
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fin…
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
Jianxu Shangguan, Jing Xu, Hang Ye +4
Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches tr…
MUSE: A Unified Agentic Harness for MLLMs
Jianglin Lu, Hailing Wang, Xu Ma +4
Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze from a screenshot or selecting t…
GazeVLA: Learning Human Intention for Robotic Manipulation
Chengyang Li, Kaiyi Xiong, Yuan Xu +3
Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works…
EgoSelf: From Memory to Personalized Egocentric Assistant
Yanshuo Wang, Yuan Xu, Xuesong Li +4
Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preference…