From the 1 of 8 linked papers with an AI index.
8 papers
TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction
Lei Jin, Yiding Ma, Xin Zhang +3
The paper introduces TacWAM, a mechanics-aware tactile world action model that predicts future tactile signals and uses them as supervision for training contact-rich robot manipula…
VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning
Shufan Zhang, Ziyue Lin, Bairun Wang +4
Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced to this field to enhance rea…
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Yu Shang, Yinzhou Tang, Yiding Ma +22
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…
Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction
Kejun Ren, Lei Jin, Tianxin Huang +2
Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across…
GraphiContact: Pose-aware Human-Scene Robust Contact Perception for Interactive Systems
Xiaojian Lin, Yaomin Shen, Junyuan Ma +7
Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In…
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
Yu Shang, Zhuohang Li, Yiding Ma +18
While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their eva…