31 citations · 57 across the 4 of their papers we have counts for
9 papers
Co-training Transformer with Videos and Images Improves Action Recognition
Bowen Zhang, Jiahui Yu, Christopher Fifty +4
In learning action recognition, models are typically pre-trained on object recognition with images, such as ImageNet, and later fine-tuned on target action recognition with videos.…
Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?
Linlu Qiu, Hexiang Hu, Bowen Zhang +2
We analyze the grounded SCAN (gSCAN) benchmark, which was recently proposed to study systematic generalization for grounded language understanding. First, we study which aspects of…
A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus
Bowen Zhang, Hexiang Hu, Joonseok Lee +5
Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…
Online Action Detection in Streaming Videos with Time Buffers
Bowen Zhang, Hao Chen, Meng Wang +1
We formulate the problem of online temporal action detection in live streaming videos, acknowledging one important property of live streaming videos that there is normally a broadc…
Learning to Represent Image and Text with Denotation Graph
Bowen Zhang, Hexiang Hu, Vihan Jain +2
Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…
Visual Storytelling via Predicting Anchor Word Embeddings in the Stories
Bowen Zhang, Hexiang Hu, Fei Sha
We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…