5 citations · 8 across the 3 of their papers we have counts for
6 papers · 1 filter
UniViTAR: Unified Vision Transformer with Native Resolution
Limeng Qiao, Yiyang Gan, Bairui Wang +4
Conventional Vision Transformer simplifies visual modeling by standardizing input resolutions, often disregarding the variability of natural visual data and compromising spatial-co…
E2E-LOAD: End-to-End Long-form Online Action Detection
Shuqiang Cao, Weixin Luo, Bairui Wang +2
Recently, there has been a growing trend toward feature-based approaches for Online Action Detection (OAD). However, these approaches have limitations due to their fixed backbone d…
Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network
Bairui Wang, Lin Ma, Wei Zhang +3
In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We const…
Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
Wei Zhang, Bairui Wang, Lin Ma +1
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…
Hierarchical Photo-Scene Encoder for Album Storytelling
Bairui Wang, Lin Ma, Wei Zhang +2
In this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two sub-enc…
Reconstruction Network for Video Captioning
Bairui Wang, Lin Ma, Wei Zhang +1
In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…