53 citations · 57 across the 2 of their papers we have counts for
3 papers
Modeling Relationships in Referential Expressions with Compositional Modular Networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas +2
People often refer to entities in an image in terms of their relationships with other entities. For example, "the black cat sitting under the table" refers to both a "black cat" en…
A Dataset for Movie Description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon +1
Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by de…
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue +3
Solving the visual symbol grounding problem has long been a goal of artificial intelligence. The field appears to be advancing closer to this goal with recent breakthroughs in deep…