9 papers
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
Motonari Kambara, Koki Seno, Tomoya Kaichi +2
We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only…
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Yusuke Takagi, Motonari Kambara, Daichi Yashima +3
In this study, we address the problem of language-guided robotic manipulation, where a robot is required to manipulate a wide range of objects based on visual observations and natu…
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
Ryosuke Takanami, Petr Khrapchenkov, Shu Morikuni +32
As robots transition from controlled settings to unstructured human environments, building generalist agents that can reliably follow natural language instructions remains a centra…
Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
Motonari Kambara, Komei Sugiura
In this work, we address the problem of predicting the future success of open-vocabulary object manipulation tasks. Conventional approaches typically determine success or failure a…
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
Kei Katsumata, Motonari Kambara, Daichi Yashima +2
We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not a…
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
Motonari Kambara, Komei Sugiura
This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based…