activity
20232026
collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

Semantic Refinement of Universal Audio Representations through Audio-Description Alignment

Lejun Min, Junyu Dai, Ruichen Zheng +7

Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of differ…

eess.AS2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan +14

Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scene…

eess.AS2024

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Xingchen Song, Chengdong Liang, Binbin Zhang +9

Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. Ho…

eess.AS2024

HydraFormer: One Encoder For All Subsampling Rates

Yaoxun Xu, Xingchen Song, Zhiyong Wu +3

In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situati…

eess.AS2023

Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition

Kaixun Huang, Ao Zhang, Binbin Zhang +3

The attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems…