72 citations · 148 across the 9 of their papers we have counts for
12 papers · 1 filter
MosaicOS: A Simple and Effective Use of Object-Centric Images for Long-Tailed Object Detection
Cheng Zhang, Tai-Yu Pan, Yandong Li +5
Many objects do not appear frequently enough in complex scenes (e.g., certain handbags in living rooms) for training an accurate object detector, but are often found frequently by…
A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
Bowen Zhang, Hexiang Hu, Joonseok Lee +5
Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…
Learning the Best Pooling Strategy for Visual Semantic Embedding
Jiacheng Chen, Hexiang Hu, Hao Wu +2
Visual Semantic Embedding (VSE) is a dominant approach for vision-language retrieval, which aims at learning a deep embedding space such that visual data are embedded close to thei…
Learning to Represent Image and Text with Denotation Graph
Bowen Zhang, Hexiang Hu, Vihan Jain +2
Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…
Visual Storytelling via Predicting Anchor Word Embeddings in the Stories
Bowen Zhang, Hexiang Hu, Fei Sha
We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…
Evaluating Text-to-Image Matching using Binary Image Selection (BISON)
Hexiang Hu, Ishan Misra, Laurens van der Maaten
Providing systems the ability to relate linguistic and visual content is one of the hallmarks of computer vision. Tasks such as text-based image retrieval and image captioning were…