28 citations · 40 across the 6 of their papers we have counts for
4 papers · 1 filter
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
Zhiwang Zhang, Dong Xu, Wanli Ouyang +1
In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where ea…
Slow Motion Matters: A Slow Motion Enhanced Network for Weakly Supervised Temporal Action Localization
Weiqi Sun, Rui Su, Qian Yu +1
Weakly supervised temporal action localization (WTAL) aims to localize actions in untrimmed videos with only weak supervision information (e.g. video-level labels). Most existing m…
Revisiting Deep Semi-supervised Learning: An Empirical Distribution Alignment Framework and Its Generalization Bound
Feiyu Wang, Qin Wang, Wen Li +2
In this work, we revisit the semi-supervised learning (SSL) problem from a new perspective of explicitly reducing empirical distribution mismatch between labeled and unlabeled samp…
Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates
Jun Liu, Amir Shahroudy, Dong Xu +2
Skeleton-based human action recognition has attracted a lot of research attention during the past few years. Recent works attempted to utilize recurrent neural networks to model th…