activity
20132026
most citedAligning where to see and what to tell: image caption with region-based attention and scene factorization

107 citations · 274 across the 28 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

cs.CV202020 cited

A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus

Bowen Zhang, Hexiang Hu, Joonseok Lee +5

Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…

cs.CL2020

AQuaMuSe: Automatically Generating Datasets for Query-Based Multi-Document Summarization

Sayali Kulkarni, Sheide Chammas, Wan Zhu +2

Summarization is the task of compressing source document(s) into coherent and succinct passages. This is a valuable tool to present users with concise and accurate sketch of the to…

cs.CV2020

Learning to Represent Image and Text with Denotation Graph

Bowen Zhang, Hexiang Hu, Vihan Jain +2

Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…

cs.LG2020

Drinking from a Firehose: Continual Learning with Web-scale Natural Language

Hexiang Hu, Ozan Sener, Fei Sha +1

Continual learning systems will interact with humans, with each other, and with the physical world through time -- and continue to learn and adapt as they do. An important open pro…

cs.AI20208 cited

BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps

Wang Zhu, Hexiang Hu, Jiacheng Chen +4

Learning to follow instructions is of fundamental importance to autonomous agents for vision-and-language navigation (VLN). In this paper, we study how an agent can navigate long p…

cs.CV20205 cited

Visual Storytelling via Predicting Anchor Word Embeddings in the Stories

Bowen Zhang, Hexiang Hu, Fei Sha

We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…