activity
20182025
most citedReconstruct and Represent Video Contents for Captioning via Reinforcement Learning

5 citations · 8 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV20251 cited

UniViTAR: Unified Vision Transformer with Native Resolution

Limeng Qiao, Yiyang Gan, Bairui Wang +4

Conventional Vision Transformer simplifies visual modeling by standardizing input resolutions, often disregarding the variability of natural visual data and compromising spatial-co…

cs.CV2023

E2E-LOAD: End-to-End Long-form Online Action Detection

Shuqiang Cao, Weixin Luo, Bairui Wang +2

Recently, there has been a growing trend toward feature-based approaches for Online Action Detection (OAD). However, these approaches have limitations due to their fixed backbone d…

cs.CV2019

Controllable Video Captioning with POS Sequence Guidance Based on Gated Fusion Network

Bairui Wang, Lin Ma, Wei Zhang +3

In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We const…

cs.CV20195 cited

Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning

Wei Zhang, Bairui Wang, Lin Ma +1

In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…

cs.CV20192 cited

Hierarchical Photo-Scene Encoder for Album Storytelling

Bairui Wang, Lin Ma, Wei Zhang +2

In this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two sub-enc…

cs.CV2018

Reconstruction Network for Video Captioning

Bairui Wang, Lin Ma, Wei Zhang +1

In this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of…