4 papers · 1 filter
Sound and Visual Representation Learning with Multiple Pretraining Tasks
Arun Balajee Vasudevan, Dengxin Dai, Luc Van Gool
Different self-supervised tasks (SSL) reveal different features from the data. The learned feature representations can exhibit different performance for each downstream task. In th…
Semantic Object Prediction and Spatial Sound Super-Resolution with Binaural Sounds
Arun Balajee Vasudevan, Dengxin Dai, Luc Van Gool
Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sou…
Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
Arun Balajee Vasudevan, Dengxin Dai, Luc Van Gool
The role of robots in society keeps expanding, bringing with it the necessity of interacting and communicating with humans. In order to keep such interaction intuitive, we provide…
Object Referring in Visual Scene with Spoken Language
Arun Balajee Vasudevan, Dengxin Dai, Luc Van Gool
Object referring has important applications, especially for human-machine interaction. While having received great attention, the task is mainly attacked with written language (tex…