8 citations · 8 across the 3 of their papers we have counts for
8 papers
Characterizing Video Question Answering with Sparsified Inputs
Shiyuan Huang, Robinson Piramuthu, Vicente Ordonez +2
In Video Question Answering, videos are often processed as a full-length sequence of frames to ensure minimal loss of information. Recent works have demonstrated evidence that spar…
CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation
Vishnu Sashank Dorbala, Gunnar Sigurdsson, Robinson Piramuthu +2
Household environments are visually diverse. Embodied agents performing Vision-and-Language Navigation (VLN) in the wild must be able to handle this diversity, while also following…
Video in 10 Bits: Few-Bit VideoQA for Efficiency and Privacy
Shiyuan Huang, Robinson Piramuthu, Shih-Fu Chang +1
In Video Question Answering (VideoQA), answering general questions about a video requires its visual information. Yet, video often contains redundant information irrelevant to the…
Visual Grounding in Video for Unsupervised Word Translation
Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh +5
There are thousands of actively spoken languages on Earth, but a single visual world. Grounding in this visual world has the potential to bridge the gap between all these languages…
Beyond the Camera: Neural Networks in World Coordinates
Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid +1
Eye movement and strategic placement of the visual field onto the retina, gives animals increased resolution of the scene and suppresses distracting information. This fundamental s…
Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid +2
In Actor and Observer we introduced a dataset linking the first and third-person video understanding domains, the Charades-Ego Dataset. In this paper we describe the egocentric asp…