251 citations · 286 across the 15 of their papers we have counts for
22 papers
Dynamic Temporal Filtering in Video Models
Fuchen Long, Zhaofan Qiu, Yingwei Pan +3
Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling powe…
SPE-Net: Boosting Point Cloud Analysis via Rotation Robustness Enhancement
Zhaofan Qiu, Yehao Li, Yu Wang +3
In this paper, we propose a novel deep architecture tailored for 3D point cloud applications, named as SPE-Net. The embedded ``Selective Position Encoding (SPE)'' procedure relies…
Explaining Cross-Domain Recognition with Interpretable Deep Classifier
Yiheng Zhang, Ting Yao, Zhaofan Qiu +1
The recent advances in deep learning predominantly construct models in their internal representations, and it is opaque to explain the rationale behind and decisions to human users…
Motion-Focused Contrastive Learning of Video Representations
Rui Li, Yiheng Zhang, Zhaofan Qiu +3
Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper…
Representing Videos as Discriminative Sub-graphs for Action Recognition
Dong Li, Zhaofan Qiu, Yingwei Pan +3
Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore…
Boosting Video Representation Learning with Multi-Faceted Integration
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo +3
Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in th…