6 papers
RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning
Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +12
Humans can infer hidden physical processes from sparse observations, yet current evaluation protocols for Vision Language Models fail to assess whether such physical reasoning is g…
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
Wenke Xia, Pei Ren, Wenbo Yu +10
Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems…
Spatial Attention: Adapting Execution Horizons for Diffusion Policies via Observation Sensitivity
Che-Sang Park, Junsu Ha, Jianlong Fu +1
Sampling action chunks via generative models has become a widely adopted methodology for robotic learning from demonstration. However, existing methods often struggle to balance re…
MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang +4
Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate this ability through offline ques…
Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention
Yanbo Mao, Jianlong Fu, Ruoxuan Zhang +2
Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We a…
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou +7
Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision…