16 citations · 91 across the 14 of their papers we have counts for
15 papers
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen +20
Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier mo…
Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning
Yixuan Pei, Zhiwu Qing, Jun Cen +6
Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to t…
Hybrid Relation Guided Set Matching for Few-shot Action Recognition
Xiang Wang, Shiwei Zhang, Zhiwu Qing +5
Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal ali…
Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +6
Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos,…
Exploring Stronger Feature for Temporal Action Localization
Zhiwu Qing, Xiang Wang, Ziyuan Huang +6
Temporal action localization aims to localize starting and ending time with action category. Limited by GPU memory, mainstream methods pre-extract features for each video. Therefor…
OadTR: Online Action Detection with Transformers
Xiang Wang, Shiwei Zhang, Zhiwu Qing +4
Most recent approaches for online action detection tend to apply Recurrent Neural Network (RNN) to capture long-range temporal structure. However, RNN suffers from non-parallelism…