3 papers
cs.LG2021
Diarisation using location tracking with agglomerative clustering
Jeremy H. M. Wong, Igor Abramovski, Xiong Xiao +1
Previous works have shown that spatial location information can be complementary to speaker embeddings for a speaker diarisation task. However, the models used often assume that sp…
cs.SD2021
Joint speaker diarisation and tracking in switching state-space model
Jeremy H. M. Wong, Yifan Gong
Speakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, an…
eess.AS2020
High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model
Jinyu Li, Rui Zhao, Eric Sun +4
While the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion…