2 papers
cs.SD2025
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
George Close, Kris Hong, Thomas Hain +1
There has been significant research effort developing neural-network-based predictors of SQ in recent years. While a primary objective has been to develop non-intrusive, i.e.~refer…
cs.MM2025
Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings
Jason Clarke, Yoshihiko Gotoh, Stefan Goetze
Audiovisual active speaker detection (ASD) addresses the task of determining the speech activity of a candidate speaker given acoustic and visual data. Typically, systems model the…