1 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
Dual-path Collaborative Generation Network for Emotional Video Captioning
Cheng Ye, Weidong Chen, Jingyu Li +2
Emotional Video Captioning is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectivel…
cs.CV2023★ 1 cited
Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation
Yichen Yuan, Yifan Wang, Lijun Wang +5
Recent leading zero-shot video object segmentation (ZVOS) works devote to integrating appearance and motion information by elaborately designing feature fusion modules and identica…
cs.CV2022★ 1 cited
Spatiotemporal Self-attention Modeling with Temporal Patch Shift for Action Recognition
Wangmeng Xiang, Chao Li, Biao Wang +3
Transformer-based methods have recently achieved great advancement on 2D image-based vision tasks. For 3D video-based tasks such as action recognition, however, directly applying s…