activity
20192022
most citedIs Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

132 citations · 138 across the 4 of their papers we have counts for

collaborators

5 papers

cs.MM20226 cited

Predict-and-Update Network: Audio-Visual Speech Recognition Inspired by Human Speech Perception

Jiadong Wang, Xinyuan Qian, Haizhou Li

Audio and visual signals complement each other in human speech perception, so do they in speech recognition. The visual hint is less evident than the acoustic hint, but more robust…

cs.SD2021

SLoClas: A Database for Joint Sound Localization and Classification

Xinyuan Qian, Bidisha Sharma, Amine El Abridi +1

In this work, we present the development of a new database, namely Sound Localization and Classification (SLoClas) corpus, for studying and analyzing sound localization and classif…

eess.AS2021132 cited

Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Ruijie Tao, Zexu Pan, Rohan Kumar Das +3

Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and lo…

cs.SD2021

Multi-target DoA Estimation with an Audio-visual Fusion Mechanism

Xinyuan Qian, Maulik Madhavi, Zexu Pan +2

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With th…

cs.SD2019

LOCATA challenge: speaker localization with a planar array

Xinyuan Qian, Andrea Cavallaro, Alessio Brutti +1

This document describes our submission to the 2018 LOCalization And TrAcking (LOCATA) challenge (Tasks 1, 3, 5). We estimate the 3D position of a speaker using the Global Coherence…