107 citations · 221 across the 17 of their papers we have counts for
11 papers · 1 filter
A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
Bowen Zhang, Hexiang Hu, Joonseok Lee +5
Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…
Learning to Represent Image and Text with Denotation Graph
Bowen Zhang, Hexiang Hu, Vihan Jain +2
Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…
Visual Storytelling via Predicting Anchor Word Embeddings in the Stories
Bowen Zhang, Hexiang Hu, Fei Sha
We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…
Classifier and Exemplar Synthesis for Zero-Shot Learning
Soravit Changpinyo, Wei-Lun Chao, Boqing Gong +1
Zero-shot learning (ZSL) enables solving a task without the need to see its examples. In this paper, we propose two ZSL frameworks that learn to synthesize parameters for novel uns…
Cross-Modal and Hierarchical Modeling of Video and Text
Bowen Zhang, Hexiang Hu, Fei Sha
Visual data and text data are composed of information at multiple granularities. A video can describe a complex scene that is composed of multiple clips or shots, where each depict…
Cross-Dataset Adaptation for Visual Question Answering
Wei-Lun Chao, Hexiang Hu, Fei Sha
We investigate the problem of cross-dataset adaptation for visual question answering (Visual QA). Our goal is to train a Visual QA model on a source dataset but apply it to another…