53 citations · 134 across the 8 of their papers we have counts for
7 papers · 1 filter
Simple Token-Level Confidence Improves Caption Correctness
Suzanne Petryk, Spencer Whitehead, Joseph E. Gonzalez +3
The ability to judge whether a caption correctly describes an image is a critical part of vision-language understanding. However, state-of-the-art models often misinterpret the cor…
Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition
Shreyank N Gowda, Marcus Rohrbach, Frank Keller +1
We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand-designed and sample the space of possible augmented dat…
FLAVA: A Foundational Language And Vision Alignment Model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami +4
State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally,…
Modeling Relationships in Referential Expressions with Compositional Modular Networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas +2
People often refer to entities in an image in terms of their relationships with other entities. For example, "the black cat sitting under the table" refers to both a "black cat" en…
Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions
Ronghang Hu, Marcus Rohrbach, Subhashini Venugopalan +1
Image segmentation from referring expressions is a joint vision and language modeling task, where the input is an image and a textual expression describing a particular region in t…
A Dataset for Movie Description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon +1
Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by de…