3 papers
eess.AS2022
Endpoint Detection for Streaming End-to-End Multi-talker ASR
Liang Lu, Jinyu Li, Yifan Gong
Streaming end-to-end multi-talker speech recognition aims at transcribing the overlapped speech from conversations or meetings with an all-neural model in a streaming fashion, whic…
cs.LG2021
Diarisation using location tracking with agglomerative clustering
Jeremy H. M. Wong, Igor Abramovski, Xiong Xiao +1
Previous works have shown that spatial location information can be complementary to speaker embeddings for a speaker diarisation task. However, the models used often assume that sp…
cs.SD2021
Joint speaker diarisation and tracking in switching state-space model
Jeremy H. M. Wong, Yifan Gong
Speakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, an…