5 papers · 1 filter
Semantic Refinement of Universal Audio Representations through Audio-Description Alignment
Lejun Min, Junyu Dai, Ruichen Zheng +7
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of differ…
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scene…
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
Xingchen Song, Chengdong Liang, Binbin Zhang +9
Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. Ho…
HydraFormer: One Encoder For All Subsampling Rates
Yaoxun Xu, Xingchen Song, Zhiyong Wu +3
In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situati…
Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition
Kaixun Huang, Ao Zhang, Binbin Zhang +3
The attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems…