Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion
arXiv:1603.09725 · doi:10.1109/TPAMI.2017.2648793
Abstract
Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants engaged in multi-party interaction while they move around and turn their heads towards the other participants rather than facing the cameras and the microphones. Multiple-person visual tracking is combined with multiple speech-source localization in order to tackle the speech-to-person association problem. The latter is solved within a novel audio-visual fusion method on the following grounds: binaural spectral features are first extracted from a microphone pair, then a supervised audio-visual alignment technique maps these features onto an image, and finally a semi-supervised clustering method assigns binaural spectral features to visible persons. The main advantage of this method over previous work is that it processes in a principled way speech signals uttered simultaneously by multiple persons. The diarization itself is cast into a latent-variable temporal graphical model that infers speaker identities and speech turns, based on the output of an audio-visual association process, executed at each time slice, and on the dynamics of the diarization variable itself. The proposed formulation yields an efficient exact inference procedure. A novel dataset, that contains audio-visual training data as well as a number of scenarios involving several participants engaged in formal and informal dialogue, is introduced. The proposed method is thoroughly tested and benchmarked with respect to several state-of-the art diarization algorithms.
14 pages, 6 figures, 5 tables
References in corpus (2)
Cited by in corpus (15)
- Multimodal Intelligence: Representation Learning, Information Fusion, and Applications
- Multimodal Machine Learning: A Survey and Taxonomy
- Tracking Gaze and Visual Focus of Attention of People Involved in Social Interaction
- Neural Network Based Reinforcement Learning for Audio-Visual Gaze Control in Human-Robot Interaction
- Online Localization and Tracking of Multiple Moving Speakers in Reverberant Environments
- UniCon: Unified Context Network for Robust Active Speaker Detection
- AVA-AVD: Audio-Visual Speaker Diarization in the Wild
- Advances in Online Audio-Visual Meeting Transcription
- ASR-Aware End-to-end Neural Diarization
- Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization
- Variational Bayesian Inference for Audio-Visual Tracking of Multiple Speakers
- Tragic Talkers: A Shakespearean Sound- and Light-Field Dataset for Audio-Visual Machine Learning Research
- Self-supervised learning for audio-visual speaker diarization
- Look Who's Talking: Active Speaker Detection in the Wild
- A cascaded multiple-speaker localization and tracking system