4 papers
Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
Jason Clarke, Yoshihiko Gotoh, Stefan Goetze
Audiovisual active speaker detection (ASD) in egocentric recordings is challenged by frequent occlusions, motion blur, and audio interference, which undermine the discernability of…
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
Jason Clarke, Yoshihiko Gotoh, Stefan Goetze
Audiovisual active speaker detection (ASD) is conventionally performed by modelling the temporal synchronisation of acoustic and visual speech cues. In egocentric recordings, howev…
Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings
Jason Clarke, Yoshihiko Gotoh, Stefan Goetze
Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the…
The impact of differences in facial features between real speakers and 3D face models on synthesized lip motions
Rabab Algadhy, Yoshihiko Gotoh, Steve Maddock
Lip motion accuracy is important for speech intelligibility, especially for users who are hard of hearing or second language learners. A high level of realism in lip movements is a…