3 papers
cs.CV2025
GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
Kei Katsumata, Yui Iioka, Naoki Hosomi +3
We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging bec…
cs.CV2023
DialMAT: Dialogue-Enabled Transformer with Moment-Based Adversarial Training
Kanta Kaneda, Ryosuke Korekata, Yuiga Wada +7
This paper focuses on the DialFRED task, which is the task of embodied instruction following in a setting where an agent can actively ask questions about the task. To address this…
cs.CV2023
Multimodal Diffusion Segmentation Model for Object Segmentation from Manipulation Instructions
Yui Iioka, Yu Yoshida, Yuiga Wada +2
In this study, we aim to develop a model that comprehends a natural language instruction (e.g., "Go to the living room and get the nearest pillow to the radio art on the wall") and…