132 citations · 195 across the 8 of their papers we have counts for
4 papers · 2 filters
A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
Bowen Zhang, Hexiang Hu, Joonseok Lee +5
Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…
Online Action Detection in Streaming Videos with Time Buffers
Bowen Zhang, Hao Chen, Meng Wang +1
We formulate the problem of online temporal action detection in live streaming videos, acknowledging one important property of live streaming videos that there is normally a broadc…
Learning to Represent Image and Text with Denotation Graph
Bowen Zhang, Hexiang Hu, Vihan Jain +2
Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…
Visual Storytelling via Predicting Anchor Word Embeddings in the Stories
Bowen Zhang, Hexiang Hu, Fei Sha
We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…