9 citations · 16 across the 7 of their papers we have counts for
7 papers
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
Yan Li, Yifei Xing, Xiangyuan Lan +3
Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising res…
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
Hao Wang, Pengzhen Ren, Zequn Jie +8
Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods…
Prompt Customization for Continual Learning
Yong Dai, Xiaopeng Hong, Yabin Wang +3
Contemporary continual learning approaches typically select prompts from a pool, which function as supplementary inputs to a pre-trained model. However, this strategy is hindered b…
Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut Features
Yi Zhang, Jitao Sang, Junyang Wang +2
Machine learning models often learn to make predictions that rely on sensitive social attributes like gender and race, which poses significant fairness risks, especially in societa…
Strip-MLP: Efficient Token Interaction for Vision MLP
Guiping Cao, Shengda Luo, Wenjian Huang +4
Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token in…
Relate auditory speech to EEG by shallow-deep attention-based network
Fan Cui, Liyong Guo, Lang He +4
Electroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDAN…