From the 1 of 27 linked papers with an AI index.
26 citations · 26 across the 8 of their papers we have counts for
12 papers · 1 filter
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
Lishan Yang, Wenxuan Song, Xi Wang +14
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…
Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers
King Hang Wong, Lingqiao Liu, Feras Dayoub
The paper investigates whether explicit joint‑torque signals can replace the implicit force cues present in leader‑follower teleoperation for transformer‑based action‑chunking poli…
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
Wenze Wang, Mehdi Hosseinzadeh, Feras Dayoub
Robotic manipulation systems that follow language instructions often execute grasp primitives in a largely single-shot manner: a model proposes an action, the robot executes it, an…
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
We present KITE, a training-free, keyframe-anchored, layout-grounded front-end that converts long robot-execution videos into compact, interpretable tokenized evidence for vision-l…
Predictive and adaptive maps for long-term visual navigation in changing environments
Lucie Halodova, Eliska Dvorakova, Filip Majer +4
In this paper, we compare different map management techniques for long-term visual navigation in changing environments. In this scenario, the navigation system needs to continuousl…
ObjectReact: Learning Object-Relative Control for Visual Navigation
Sourav Garg, Dustin Craggs, Vineeth Bhat +5
Visual navigation using only a single camera and a topological map has recently become an appealing alternative to methods that require additional sensors and 3D maps. This is typi…