24 citations · 24 across the 2 of their papers we have counts for
2 papers
cs.SD2026
Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification
Peng Jia, Li Dai, Jia Li +3
Accurate and robust multimodal speaker identification is essential for multimedia understanding and biometric authentication. However, real-world polyglot scenarios pose two key ch…
cs.CV2023★ 24 cited
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
Zijie Song, Zhenzhen Hu, Yuanen Zhou +3
Cross-lingual image captioning is a challenging task that requires addressing both cross-lingual and cross-modal obstacles in multimedia analysis. The crucial issue in this task is…