most citedDM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.RO20241 cited

DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions

Ryosuke Korekata, Kanta Kaneda, Shunya Nagashima +2

In this study, we aim to develop a domestic service robot (DSR) that, guided by open-vocabulary instructions, can carry everyday objects to the specified pieces of furniture. Few e…

cs.CV2024

Polos: Multimodal Metric Learning from Human Feedback for Image Captioning

Yuiga Wada, Kanta Kaneda, Daichi Saito +1

Establishing an automatic evaluation metric that closely aligns with human judgments is essential for effectively developing image captioning models. Recent data-driven metrics hav…

cs.RO2023

Learning-To-Rank Approach for Identifying Everyday Objects Using a Physical-World Search Engine

Kanta Kaneda, Shunya Nagashima, Ryosuke Korekata +2

Domestic service robots offer a solution to the increasing demand for daily care and support. A human-in-the-loop approach that combines automation and operator intervention is con…

cs.CV2023

DialMAT: Dialogue-Enabled Transformer with Moment-Based Adversarial Training

Kanta Kaneda, Ryosuke Korekata, Yuiga Wada +7

This paper focuses on the DialFRED task, which is the task of embodied instruction following in a setting where an agent can actively ask questions about the task. To address this…

cs.CV2023

JaSPICE: Automatic Evaluation Metric Using Predicate-Argument Structures for Image Captioning Models

Yuiga Wada, Kanta Kaneda, Komei Sugiura

Image captioning studies heavily rely on automatic evaluation metrics such as BLEU and METEOR. However, such n-gram-based metrics have been shown to correlate poorly with human eva…