most citedDetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

2 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20232 cited

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-training via Word-Region Alignment

Lewei Yao, Jianhua Han, Xiaodan Liang +4

This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike…

cs.CV20231 cited

Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection, Segmentation, and Depth Estimation

Hanrong Ye, Dan Xu

This report serves as a supplementary document for TaskPrompter, detailing its implementation on a new joint 2D-3D multi-task learning benchmark based on Cityscapes-3D. TaskPrompte…

cs.CV20232 cited

You Only Train Once: Multi-Identity Free-Viewpoint Neural Human Rendering from Monocular Videos

Jaehyeok Kim, Dongyoon Wee, Dan Xu

We introduce You Only Train Once (YOTO), a dynamic human generation framework, which performs free-viewpoint rendering of different human identities with distinct motions, via only…

cs.LG2022

Exploring Adversarial Examples and Adversarial Robustness of Convolutional Neural Networks by Mutual Information

Jiebao Zhang, Wenhua Qian, Rencan Nie +2

A counter-intuitive property of convolutional neural networks (CNNs) is their inherent susceptibility to adversarial examples, which severely hinders the application of CNNs in sec…

cs.CV2022

Network Binarization via Contrastive Learning

Yuzhang Shang, Dan Xu, Ziliang Zong +2

Neural network binarization accelerates deep models by quantizing their weights and activations into 1-bit. However, there is still a huge performance gap between Binary Neural Net…