133 citations · 268 across the 20 of their papers we have counts for
22 papers · 1 filter
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
Zhili Liu, Kai Chen, Jianhua Han +4
Masked Autoencoder~(MAE) is a prevailing self-supervised learning method that achieves promising results in model pre-training. However, when the various downstream tasks have data…
Mixed Autoencoder for Self-supervised Visual Representation Learning
Kai Chen, Zhili Liu, Lanqing Hong +3
Masked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation s…
Generative Negative Text Replay for Continual Vision-Language Pretraining
Shipeng Yan, Lanqing Hong, Hang Xu +4
Vision-language pre-training (VLP) has attracted increasing attention recently. With a large amount of image-text pairs, VLP models trained with contrastive loss have achieved impr…
DevNet: Self-supervised Monocular Depth Learning via Density Volume Construction
Kaichen Zhou, Lanqing Hong, Changhao Chen +4
Self-supervised depth learning from monocular images normally relies on the 2D pixel-wise photometric relation between temporally adjacent image frames. However, they neither fully…
One Million Scenes for Autonomous Driving: ONCE Dataset
Jiageng Mao, Minzhe Niu, Chenhan Jiang +10
Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On th…
Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
Lewei Yao, Renjie Pi, Hang Xu +3
We propose Joint-DetNAS, a unified NAS framework for object detection, which integrates 3 key components: Neural Architecture Search, pruning, and Knowledge Distillation. Instead o…