activity
20162023
most citedCo-training Transformer with Videos and Images Improves Action Recognition

31 citations · 60 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV20233 cited

Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts

Erik Daxberger, Floris Weers, Bowen Zhang +7

Sparse Mixture-of-Experts models (MoEs) have recently gained popularity due to their ability to decouple model size from inference efficiency by only activating a small subset of t…

cs.CV202131 cited

Co-training Transformer with Videos and Images Improves Action Recognition

Bowen Zhang, Jiahui Yu, Christopher Fifty +4

In learning action recognition, models are typically pre-trained on object recognition with images, such as ImageNet, and later fine-tuned on target action recognition with videos.…

cs.CV202020 cited

A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus

Bowen Zhang, Hexiang Hu, Joonseok Lee +5

Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, b…

cs.CV2020

Online Action Detection in Streaming Videos with Time Buffers

Bowen Zhang, Hao Chen, Meng Wang +1

We formulate the problem of online temporal action detection in live streaming videos, acknowledging one important property of live streaming videos that there is normally a broadc…

cs.CV2020

Learning to Represent Image and Text with Denotation Graph

Bowen Zhang, Hexiang Hu, Vihan Jain +2

Learning to fuse vision and language information and representing them is an important research problem with many applications. Recent progresses have leveraged the ideas of pre-tr…

cs.CV20205 cited

Visual Storytelling via Predicting Anchor Word Embeddings in the Stories

Bowen Zhang, Hexiang Hu, Fei Sha

We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jo…