11 papers
Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations
Simon Holk, Ryosuke Takanami, Tatsuya Matsushima +4
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with…
YUBI: Yielding Universal Bidigital Interface for Bimanual Dexterous Manipulation at Scale
Takehiko Ohkawa, Jumpei Arima, Yuki Noguchi +16
We introduce Yielding Universal Bidigital Interface (YUBI), a finger-aligned gripper designed to enable intuitive, ergonomic, and scalable data collection for bimanual dexterous ma…
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often inf…
Continuous Reasoning for Vision-Language-Action
Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals o…
Touch2Insert: Zero-Shot Peg Insertion by Touching Intersections of Peg and Hole
Masaru Yajima, Yuma Shin, Rei Kawakami +2
Reliable insertion of industrial connectors remains a central challenge in robotics, requiring sub-millimeter precision under uncertainty and often without full visual access. Visi…
Simultaneous Extrinsic Contact and In-Hand Pose Estimation via Distributed Tactile Sensing
Mark Van der Merwe, Kei Ota, Dmitry Berenson +2
Prehensile autonomous manipulation, such as peg insertion, tool use, or assembly, require precise in-hand understanding of the object pose and the extrinsic contacts made during in…