16 citations · 110 across the 18 of their papers we have counts for
18 papers · 1 filter
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen +20
Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier mo…
Replace Anyone in Videos
Xiang Wang, Shiwei Zhang, Haonan Qiu +7
The field of controllable human-centric video generation has witnessed remarkable progress, particularly with the advent of diffusion models. However, achieving precise and localiz…
Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning
Yixuan Pei, Zhiwu Qing, Jun Cen +6
Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to t…
Benchmarking Unsupervised Anomaly Detection and Localization
Ye Zheng, Xiang Wang, Yu Qi +2
Unsupervised anomaly detection and localization, as of one the most practical and challenging problems in computer vision, has received great attention in recent years. From the ti…
Hybrid Relation Guided Set Matching for Few-shot Action Recognition
Xiang Wang, Shiwei Zhang, Zhiwu Qing +5
Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal ali…
Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +6
Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos,…