132 citations · 138 across the 4 of their papers we have counts for
5 papers
Predict-and-Update Network: Audio-Visual Speech Recognition Inspired by Human Speech Perception
Jiadong Wang, Xinyuan Qian, Haizhou Li
Audio and visual signals complement each other in human speech perception, so do they in speech recognition. The visual hint is less evident than the acoustic hint, but more robust…
SLoClas: A Database for Joint Sound Localization and Classification
Xinyuan Qian, Bidisha Sharma, Amine El Abridi +1
In this work, we present the development of a new database, namely Sound Localization and Classification (SLoClas) corpus, for studying and analyzing sound localization and classif…
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das +3
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and lo…
Multi-target DoA Estimation with an Audio-visual Fusion Mechanism
Xinyuan Qian, Maulik Madhavi, Zexu Pan +2
Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With th…
LOCATA challenge: speaker localization with a planar array
Xinyuan Qian, Andrea Cavallaro, Alessio Brutti +1
This document describes our submission to the 2018 LOCalization And TrAcking (LOCATA) challenge (Tasks 1, 3, 5). We estimate the 3D position of a speaker using the Global Coherence…