3 papers
cs.LG2026
From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation
Shuchang Ye, Jinqiang Yu, Zhujun Xiao +6
Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation t…
cs.LG2024
A Comprehensive Solution to Connect Speech Encoder and Large Language Model for ASR
Van Tung Pham, Yist Lin, Tao Han +4
Recent works have shown promising results in connecting speech encoders to large language models (LLMs) for speech recognition. However, several limitations persist, including limi…
eess.AS2024
A Large-Scale Evaluation of Speech Foundation Models
Shu-wen Yang, Heng-Jui Chang, Zili Huang +18
The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling a…