works on

From the 2 of 7 linked papers with an AI index.

collaborators

7 papers

cs.MM2026

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +12

The paper presents RetroHolmes, a benchmark for testing vision‑language models on retrospective physical process reasoning—inferring hidden causes from image outcomes—and shows tha…

cs.AI2026

MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

Ruoxuan Zhang, Qiaoqiao Wan, Zhengguang Wang +4

The paper introduces MindClaw, a closed-loop framework that lets embodied agents reason about human mental states in real time and intervene only when assistance is needed.

cs.RO2026

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

Wenke Xia, Pei Ren, Wenbo Yu +10

Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems…

cs.RO2026

Latent Policy Steering through One-Step Flow Policies

Hokyun Im, Andrey Kolobov, Jianlong Fu +1

Offline reinforcement learning (RL) allows robots to learn from offline datasets without risky exploration. Yet, offline RL's performance often hinges on a brittle trade-off betwee…

cs.RO2026

Spatial Attention: Adapting Execution Horizons for Diffusion Policies via Observation Sensitivity

Che-Sang Park, Junsu Ha, Jianlong Fu +1

Sampling action chunks via generative models has become a widely adopted methodology for robotic learning from demonstration. However, existing methods often struggle to balance re…

cs.AI2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou +7

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision…