activity
20182022
most citedPointDAN: A Multi-Scale 3D Domain Adaption Network for Point Cloud Representation

74 citations · 142 across the 4 of their papers we have counts for

collaborators

12 papers

cs.CV20228 cited

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +8

Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…

cs.LG202150 cited

Graph-MLP: Node Classification without Message Passing in Graph

Yang Hu, Haoxuan You, Zhecan Wang +3

Graph Neural Network (GNN) has been demonstrated its effectiveness in dealing with non-Euclidean structural data. Both spatial-based and spectral-based GNNs are relying on adjacenc…

cs.CL2020

Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions

Liunian Harold Li, Haoxuan You, Zhecan Wang +3

Pre-trained contextual vision-and-language (V&L) models have achieved impressive performance on various benchmarks. However, existing models require a large amount of parallel imag…

cs.CV202010 cited

Learning Visual Commonsense for Robust Scene Graph Generation

Alireza Zareian, Zhecan Wang, Haoxuan You +1

Scene graph generation models understand the scene through object and predicate recognition, but are prone to mistakes due to the challenges of perception in the wild. Perception e…

cs.CV201974 cited

PointDAN: A Multi-Scale 3D Domain Adaption Network for Point Cloud Representation

Can Qin, Haoxuan You, Lichen Wang +2

Domain Adaptation (DA) approaches achieved significant improvements in a wide range of machine learning and computer vision tasks (i.e., classification, detection, and segmentation…

cs.CV2019

Multi-modality Latent Interaction Network for Visual Question Answering

Peng Gao, Haoxuan You, Zhanpeng Zhang +2

Exploiting relationships between visual regions and question words have achieved great success in learning multi-modality features for Visual Question Answering (VQA). However, we…