3 papers
cs.CV2025
AMMKD: Adaptive Multimodal Multi-teacher Distillation for Lightweight Vision-Language Models
Yuqi Li, Chuanguang Yang, Junhao Dong +6
The success of large-scale visual language pretraining (VLP) models has driven widespread adoption of image-text retrieval tasks. However, their deployment on mobile devices remain…
cs.CL2025
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
Tianyi Xu, Hongjie Chen, Wang Qing +6
Large-scale training corpora have significantly improved the performance of ASR models. Unfortunately, due to the relative scarcity of data, Chinese accents and dialects remain a c…
cs.SD2025
SepPrune: Structured Pruning for Efficient Deep Speech Separation
Yuqi Li, Kai Li, Xin Yin +6
Although deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational…