activity
20202022
most citedTAda! Temporally-Adaptive Convolutions for Video Understanding

31 citations · 173 across the 21 of their papers we have counts for

collaborators
Showing 2022Show all

6 papers · 1 filter

cs.CV2022★ 12 cited

Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning

Yixuan Pei, Zhiwu Qing, Jun Cen +6

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to t…

cs.LG2022★ 14 cited

Grow and Merge: A Unified Framework for Continuous Categories Discovery

Xinwei Zhang, Jianwen Jiang, Yutong Feng +6

Although a number of studies are devoted to novel category discovery, most of them assume a static setting where both labeled and unlabeled data are given at once for finding new c…

cs.CV2022★ 29 cited

RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

Hangjie Yuan, Jianwen Jiang, Samuel Albanie +4

The task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications. Prior…

cs.CV2022★ 1 cited

Open-world Semantic Segmentation for LIDAR Point Clouds

Jun Cen, Peng Yun, Shiwei Zhang +5

Current methods for LIDAR semantic segmentation are not robust enough for real-world applications, e.g., autonomous driving, since it is closed-set and static. The closed-set assum…

cs.CV2022★ 7 cited

Hybrid Relation Guided Set Matching for Few-shot Action Recognition

Xiang Wang, Shiwei Zhang, Zhiwu Qing +5

Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal ali…

cs.CV2022

Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency

Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +6

Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos,…