From the 1 of 6 linked papers with an AI index.
6 papers
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
Lishan Yang, Wenxuan Song, Xi Wang +14
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…
Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers
King Hang Wong, Lingqiao Liu, Feras Dayoub
The paper investigates whether explicit joint‑torque signals can replace the implicit force cues present in leader‑follower teleoperation for transformer‑based action‑chunking poli…
TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models
Wenbo Zhang, Jianxiong Li, Shuai Yang +4
Vision-Language-Action (VLA) models trained on large-scale data have made remarkable progress, but they remain vulnerable to distribution shifts at deployment time. Recent VLA mode…
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
Wenbo Zhang, Tianrun Hu, Hanbo Zhang +7
We present Chain-of-Action (CoA), a novel visuo-motor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s)…
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models
Ankit Yadav, Lingqiao Liu, Yuankai Qi
This work investigates the capabilities of current vision-language models (VLMs) in visual understanding and attribute measurement of primitive shapes using a benchmark focused on…
Embodied Domain Adaptation for Object Detection
Xiangyu Shi, Yanyuan Qiao, Lingqiao Liu +1
Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects an…