activity
20182026
most citedAdvances in Online Audio-Visual Meeting Transcription

17 citations · 33 across the 6 of their papers we have counts for

collaborators

12 papers

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

eess.AS20221 cited

Target Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization

Dongmei Wang, Xiong Xiao, Naoyuki Kanda +2

This paper describes a speaker diarization model based on target speaker voice activity detection (TS-VAD) using transformers. To overcome the original TS-VAD model's drawback of b…

eess.AS20211 cited

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

Naoyuki Kanda, Xiong Xiao, Jian Wu +6

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules…

eess.AS202112 cited

Speaker attribution with voice profiles by graph-based semi-supervised learning

Jixuan Wang, Xiong Xiao, Jian Wu +3

Speaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice prof…

eess.AS2020

Microsoft Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2020

Xiong Xiao, Naoyuki Kanda, Zhuo Chen +10

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recogniti…

eess.AS2020

Speaker diarization with session-level speaker embedding refinement using graph neural networks

Jixuan Wang, Xiong Xiao, Jian Wu +3

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a gl…