1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2024
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
Yihan Wu, Soumi Maiti, Yifan Peng +6
Recent advancements in language models have significantly enhanced performance in multiple speech-related tasks. Existing speech language models typically utilize task-dependent pr…
eess.AS2023★ 1 cited
Adapting Multi-Lingual ASR Models for Handling Multiple Talkers
Chenda Li, Yao Qian, Zhuo Chen +5
State-of-the-art large-scale universal speech models (USMs) show a decent automatic speech recognition (ASR) performance across multiple domains and languages. However, it remains…
eess.AS2023
Target Sound Extraction with Variable Cross-modality Clues
Chenda Li, Yao Qian, Zhuo Chen +5
Automatic target sound extraction (TSE) is a machine learning approach to mimic the human auditory perception capability of attending to a sound source of interest from a mixture o…