90 citations · 183 across the 10 of their papers we have counts for
4 papers · 1 filter
Exploring Robust Face-Voice Matching in Multilingual Environments
Jiehui Tang, Xiaofei Wang, Zhen Xiao +3
This paper presents Team Xaiofei's innovative approach to exploring Face-Voice Association in Multilingual Environments (FAME) at ACM Multimedia 2024. We focus on the impact of dif…
Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection
Xingyu Peng, Yan Bai, Chen Gao +5
Open-Vocabulary Detection (OVD) is the task of detecting all interesting objects in a given scene without predefined object classes. Extensive work has been done to deal with the O…
Domain Game: Disentangle Anatomical Feature for Single Domain Generalized Segmentation
Hao Chen, Hongrun Zhang, U Wang Chan +3
Single domain generalization aims to address the challenge of out-of-distribution generalization problem with only one source domain available. Feature distanglement is a classic s…
Eliminating Cross-modal Conflicts in BEV Space for LiDAR-Camera 3D Object Detection
Jiahui Fu, Chen Gao, Zitian Wang +4
Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird's-eye view (BEV) representation space. However, our empirical findi…