16 papers
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
Simon Holk, Ryosuke Takanami, Tatsuya Matsushima +4
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with…
FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning
Takanori Yoshimoto, Yang Hu, Naruya Kondo +1
Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed…
YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale
Takehiko Ohkawa, Jumpei Arima, Yuki Noguchi +16
We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous ma…
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often inf…
Continuous Reasoning for Vision-Language-Action
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals o…
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni +32
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a centra…