3 citations · 7 across the 6 of their papers we have counts for
6 papers
Lighten CARAFE: Dynamic Lightweight Upsampling with Guided Reassemble Kernels
Ruigang Fu, Qingyong Hu, Xiaohu Dong +3
As a fundamental operation in modern machine vision models, feature upsampling has been widely used and investigated in the literatures. An ideal upsampling operation should be lig…
Divide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries
Qi Song, Qingyong Hu, Chi Zhang +2
3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that…
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
Hao Lu, Xuesong Niu, Jiyao Wang +12
Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in l…
FedGT: Federated Node Classification with Scalable Graph Transformer
Zaixi Zhang, Qingyong Hu, Yang Yu +2
Graphs are widely used to model relational data. As graphs are getting larger and larger in real-world scenarios, there is a trend to store and compute subgraphs in multiple local…
PVP: Pre-trained Visual Parameter-Efficient Tuning
Zhao Song, Ke Yang, Naiyang Guan +3
Large-scale pre-trained transformers have demonstrated remarkable success in various computer vision tasks. However, it is still highly challenging to fully fine-tune these models…
Backdoor Defense via Deconfounded Representation Learning
Zaixi Zhang, Qi Liu, Zhicai Wang +2
Deep neural networks (DNNs) are recently shown to be vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by injecting a few poisoned examples in…