most citedSpeaker diarisation using 2D self-attentive combination of embeddings

2 citations · 2 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2023

It HAS to be Subjective: Human Annotator Simulation via Zero-shot Density Estimation

Wen Wu, Wenlin Chen, Chao Zhang +1

Human annotator simulation (HAS) serves as a cost-effective substitute for human evaluation such as data annotation and system assessment. Human perception and behaviour during hum…

cs.CL2023

Can Contextual Biasing Remain Effective with Whisper and GPT-2?

Guangzhi Sun, Xianrui Zheng, Chao Zhang +1

End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the larg…

cs.CL20231 cited

Graph Neural Networks for Contextual ASR with the Tree-Constrained Pointer Generator

Guangzhi Sun, Chao Zhang, Phil Woodland

The incorporation of biasing words obtained through contextual knowledge is of paramount importance in automatic speech recognition (ASR) applications. This paper proposes an innov…

cs.CL2022

Turn-Taking Prediction for Natural Conversational Speech

Shuo-yiin Chang, Bo Li, Tara N. Sainath +4

While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice qu…

cs.CL20192 cited

Speaker diarisation using 2D self-attentive combination of embeddings

Guangzhi Sun, Chao Zhang, Phil Woodland

Speaker diarisation systems often cluster audio segments using speaker embeddings such as i-vectors and d-vectors. Since different types of embeddings are often complementary, this…