8 papers
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR
Yuan Xie, Jiaqi Song, Xianliang Wang +3
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achiev…
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
Yuan Xie, Jiaqi Song, Guang Qiu +9
Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrat…
InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark
Shiyu Wang, Ziyu Liu, Chaoyi Yu +6
Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benc…
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
Zhengjia Zhong, Shuyan Ke, Zaizhou Lin +3
Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under…
ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception
Huanzhen Wang, Ziheng Zhou, Jiaqi Song +4
Dynamic facial expression recognition in the wild remains challenging due to data scarcity and long-tail distributions, which hinder models from effectively learning the temporal d…
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
Yuan Xie, Jiaqi Song, Guang Qiu +4
Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performan…