19 papers
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
Chenghao Gu, Hanyang Yu, Jingbo Zhang +7
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse rea…
Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models
Jingyan Jiang, Yaru Sun, Xiao Chen +5
Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unr…
TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding
Kepeng Yang, Dongxuan Liu, Rongxin Gao +8
Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it v…
ChainVLA: Chaining Vision-Language-Action Queries through a Unified Execution State for Long-Horizon Manipulation
Yuzhi Huang, Weijue Bu, Ziyi Xiong +4
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunke…
DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs
Rongxin Gao, Yuzhi Huang, Dongxuan Liu +8
4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling embodied interaction. While c…
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
Yuzhi Huang, Jie Wu, Weijue Bu +9
Enabling reliable long-horizon robotic manipulation is a crucial step toward open-world embodied intelligence. However, VLM-based planners treat each step as an isolated observatio…