103 citations · 121 across the 82 of their papers we have counts for
18 papers · 1 filter
Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision
Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima +2
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action mo…
MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?
Kohei Sendai, Tatsuya Matsushima, Yusuke Iwasawa
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains un…
DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation
Makoto Sato, Tatsuya Matsushima, Yutaka Matsuo +1
Vision-language-action (VLA) models have made strong progress in language-conditioned robot manipulation, but improving their performance in a new workspace still often requires ac…
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
Simon Holk, Ryosuke Takanami, Tatsuya Matsushima +4
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with…
Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems
Zhongchao Zhou, Yixuan Xie, Wenwei Yu +5
Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchm…
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change…