1 paper · 1 filter
Kei Katsumata, Motonari Kambara, Daichi Yashima +2
We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not a…