16 citations · 36 across the 10 of their papers we have counts for
7 papers · 1 filter
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
Yuhang Song, Mario Gianni, Chenguang Yang +4
This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural languag…
MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling
Diwei Huang, Kunyang Lin, Peihao Chen +2
Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-s…
Event-Guided Procedure Planning from Instructional Videos with Text Supervision
An-Lan Wang, Kun-Yu Lin, Jia-Run Du +2
In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial…
Learning Vision-and-Language Navigation from YouTube Videos
Kunyang Lin, Peihao Chen, Diwei Huang +3
Vision-and-language navigation (VLN) requires an embodied agent to navigate in realistic 3D environments using natural language instructions. Existing VLN methods suffer from train…
DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition
Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin +4
As a de facto solution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive f…
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin +4
We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions oft…