activity
20172022
most citedLearning Spatio-Temporal Representation with Pseudo-3D Residual Networks

251 citations · 286 across the 15 of their papers we have counts for

collaborators

22 papers

cs.CV2022

Dynamic Temporal Filtering in Video Models

Fuchen Long, Zhaofan Qiu, Yingwei Pan +3

Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling powe…

cs.CV2022

SPE-Net: Boosting Point Cloud Analysis via Rotation Robustness Enhancement

Zhaofan Qiu, Yehao Li, Yu Wang +3

In this paper, we propose a novel deep architecture tailored for 3D point cloud applications, named as SPE-Net. The embedded ``Selective Position Encoding (SPE)'' procedure relies…

cs.CV20221 cited

Explaining Cross-Domain Recognition with Interpretable Deep Classifier

Yiheng Zhang, Ting Yao, Zhaofan Qiu +1

The recent advances in deep learning predominantly construct models in their internal representations, and it is opaque to explain the rationale behind and decisions to human users…

cs.CV20224 cited

Motion-Focused Contrastive Learning of Video Representations

Rui Li, Yiheng Zhang, Zhaofan Qiu +3

Motion, as the most distinct phenomenon in a video to involve the changes over time, has been unique and critical to the development of video representation learning. In this paper…

cs.CV20223 cited

Representing Videos as Discriminative Sub-graphs for Action Recognition

Dong Li, Zhaofan Qiu, Yingwei Pan +3

Human actions are typically of combinatorial structures or patterns, i.e., subjects, objects, plus spatio-temporal interactions in between. Discovering such structures is therefore…

cs.CV2022

Boosting Video Representation Learning with Multi-Faceted Integration

Zhaofan Qiu, Ting Yao, Chong-Wah Ngo +3

Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in th…