10 citations · 14 across the 8 of their papers we have counts for
5 papers
Speech-based Slot Filling using Large Language Models
Guangzhi Sun, Shutong Feng, Dongcheng Jiang +3
Recently, advancements in large language models (LLMs) have shown an unprecedented ability across various language tasks. This paper investigates the potential application of LLMs…
TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
Jeff Hwang, Moto Hira, Caroline Chen +21
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing…
Conditional Diffusion Model for Target Speaker Extraction
Theodor Nguyen, Guangzhi Sun, Xianrui Zheng +2
We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuou…
Enhancing Quantised End-to-End ASR Models via Personalisation
Qiuming Zhao, Guangzhi Sun, Chao Zhang +2
Recent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Mo…
Can Contextual Biasing Remain Effective with Whisper and GPT-2?
Guangzhi Sun, Xianrui Zheng, Chao Zhang +1
End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the larg…