From the 1 of 18 linked papers with an AI index.
18 papers
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
Lishan Yang, Wenxuan Song, Xi Wang +14
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…
Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers
King Hang Wong, Lingqiao Liu, Feras Dayoub
The paper investigates whether explicit joint‑torque signals can replace the implicit force cues present in leader‑follower teleoperation for transformer‑based action‑chunking poli…
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
Bo Miao, Weijia Liu, Jun Luo +8
Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-le…
Detecting Precise Hand Touch Moments in Egocentric Video
Huy Anh Nguyen, Feras Dayoub, Minh Hoai
We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is crucial for augmented reali…
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
Wenze Wang, Mehdi Hosseinzadeh, Feras Dayoub
Robotic manipulation systems that follow language instructions often execute grasp primitives in a largely single-shot manner: a model proposes an action, the robot executes it, an…
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
We present KITE, a training-free, keyframe-anchored, layout-grounded front-end that converts long robot-execution videos into compact, interpretable tokenized evidence for vision-l…