activity
20172024
most citedUniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

108 citations · 149 across the 11 of their papers we have counts for

collaborators

17 papers

cs.CV2024

Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models

Xiaoshi Wu, Yiming Hao, Manyuan Zhang +5

Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an a…

cs.CV2022

Teach-DETR: Better Training DETR with Teachers

Linjiang Huang, Kaixin Lu, Guanglu Song +4

In this paper, we present a novel training scheme, namely Teach-DETR, to learn better DETR-based detectors from versatile teacher detectors. We show that the predicted boxes from t…

cs.CV20222 cited

Large-batch Optimization for Dense Visual Predictions

Zeyue Xue, Jianming Liang, Guanglu Song +4

Training a large-scale deep neural network in a large-scale dataset is challenging and time-consuming. The recent breakthrough of large-batch optimization is a promising way to tac…

cs.CV2022

Towards Robust Face Recognition with Comprehensive Search

Manyuan Zhang, Guanglu Song, Yu Liu +1

Data cleaning, architecture, and loss function design are important factors contributing to high-performance face recognition. Previously, the research community tries to improve t…

cs.CV2022

Unifying Visual Perception by Dispersible Points Learning

Jianming Liang, Guanglu Song, Biao Leng +1

We present a conceptually simple, flexible, and universal visual perception head for variant visual tasks, e.g., classification, object detection, instance segmentation and pose es…

cs.CV2022108 cited

UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Kunchang Li, Yali Wang, Peng Gao +4

It is a challenging task to learn rich and multi-scale spatiotemporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between vi…