36 citations · 74 across the 3 of their papers we have counts for
3 papers
FLAVA: A Foundational Language And Vision Alignment Model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami +4
State-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally,…
Modeling Relationships in Referential Expressions with Compositional Modular Networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas +2
People often refer to entities in an image in terms of their relationships with other entities. For example, "the black cat sitting under the table" refers to both a "black cat" en…
Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions
Ronghang Hu, Marcus Rohrbach, Subhashini Venugopalan +1
Image segmentation from referring expressions is a joint vision and language modeling task, where the input is an image and a textual expression describing a particular region in t…