325 citations · 1.4k across the 92 of their papers we have counts for
38 papers · 1 filter
Point-Query Quadtree for Crowd Counting, Localization, and More
Chengxin Liu, Hao Lu, Zhiguo Cao +1
We show that crowd counting can be viewed as a decomposable point querying process. This formulation enables arbitrary points as input and jointly reasons whether the points are cr…
ALIP: Adaptive Language-Image Pre-training with Synthetic Caption
Kaicheng Yang, Jiankang Deng, Xiang An +5
Contrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected…
PNT-Edge: Towards Robust Edge Detection with Noisy Labels by Learning Pixel-level Noise Transitions
Wenjie Xuan, Shanshan Zhao, Yu Yao +5
Relying on large-scale training data with pixel-level labels, previous edge detection methods have achieved high performance. However, it is hard to manually label edges accurately…
Why do CNNs excel at feature extraction? A mathematical explanation
Vinoth Nandakumar, Arush Tagade, Tongliang Liu
Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification b…
Towards Label-free Scene Understanding by Vision Foundation Models
Runnan Chen, Youquan Liu, Lingdong Kong +5
Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classificati…
Private Gradient Estimation is Useful for Generative Modeling
Bochao Liu, Pengju Wang, Weijia Guo +4
While generative models have proved successful in many domains, they may pose a privacy leakage risk in practical deployment. To address this issue, differentially private generati…