1 paper
Kei Katsumata, Yui Iioka, Naoki Hosomi +3
We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging bec…