1 paper
Dongwoo Kang, Akhil Perincherry, Zachary Coalson +3
An emerging paradigm in vision-and-language navigation (VLN) is the use of history-aware multi-modal transformer models. Given a language instruction, these models process observat…