30 citations · 38 across the 19 of their papers we have counts for
19 papers
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
Motonari Kambara, Komei Sugiura
This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based…
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
Kazuki Matsuda, Yuiga Wada, Komei Sugiura
In this work, we address the challenge of developing automatic evaluation metrics for image captioning, with a particular focus on robustness against hallucinations. Existing metri…
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
Miyu Goko, Motonari Kambara, Daichi Saito +2
In this study, we consider the problem of predicting task success for open-vocabulary manipulation by a manipulator, based on instruction sentences and egocentric images before and…
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
Ryosuke Korekata, Kanta Kaneda, Shunya Nagashima +2
In this study, we aim to develop a domestic service robot (DSR) that, guided by open-vocabulary instructions, can carry everyday objects to the specified pieces of furniture. Few e…
Moment-based Adversarial Training for Embodied Language Comprehension
Shintaro Ishikawa, Komei Sugiura
In this paper, we focus on a vision-and-language task in which a robot is instructed to execute household tasks. Given an instruction such as "Rinse off a mug and place it in the c…
Target-dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots
Shintaro Ishikawa, Komei Sugiura
Currently, domestic service robots have an insufficient ability to interact naturally through language. This is because understanding human instructions is complicated by various a…