11 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CL2022★ 11 cited
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities
Zhongzhi Chen, Guang Liu, Bo-Wen Zhang +3
In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimod…
cs.SD2022★ 2 cited
Phonemic Adversarial Attack against Audio Recognition in Real World
Jiakai Wang, Zhendong Chen, Zixin Yin +2
Recently, adversarial attacks for audio recognition have attracted much attention. However, most of the existing studies mainly rely on the coarse-grain audio features at the insta…