20 citations · 40 across the 18 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 3 cited
Mind the Gap: Improving Success Rate of Vision-and-Language Navigation by Revisiting Oracle Success Routes
Chongyang Zhao, Yuankai Qi, Qi Wu
Vision-and-Language Navigation (VLN) aims to navigate to the target location by following a given instruction. Unlike existing methods focused on predicting a more accurate action…
cs.CV2021★ 3 cited
LocFormer: Enabling Transformers to Perform Temporal Moment Localization on Long Untrimmed Videos With a Feature Sampling Approach
Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando +2
We propose LocFormer, a Transformer-based model for video grounding which operates at a constant memory footprint regardless of the video length, i.e. number of frames. LocFormer i…