12 citations · 15 across the 8 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios
Bingshen Mu, Xian Shi, Xiong Wang +6
Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who s…
eess.AS2026
Qwen-Audio-VAE Technical Report
Ziyue Jiang, Dake Guo, Zekai Zhang +11
We introduce \textbf{Qwen-Audio-VAE}, a suite of low-bitrate, fast-encoding continuous audio autoencoders designed for scalable general audio generation. The model is built around…
eess.AS2025
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
He Wang, Linhan Ma, Dake Guo +4
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…