3 papers
eess.AS2024
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models
Quan Wang, Yiling Huang, Guanlong Zhao +3
In this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system. Various goals can be a…
eess.AS2023
Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network
Yiling Huang, Weiran Wang, Guanlong Zhao +3
While standard speaker diarization attempts to answer the question "who spoken when", most of relevant applications in reality are more interested in determining "who spoken what".…
eess.AS2023
USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models
Guanlong Zhao, Yongqiang Wang, Jason Pelecanos +5
We introduce a multilingual speaker change detection model (USM-SCD) that can simultaneously detect speaker turns and perform ASR for 96 languages. This model is adapted from a spe…