6 papers
How To Train Your Deep Multi-Object Tracker
Yihong Xu, Aljosa Osep, Yutong Ban +3
The recent trend in vision-based multi-object tracking (MOT) is heading towards leveraging the representational power of deep learning to jointly learn to detect and track objects.…
A cascaded multiple-speaker localization and tracking system
Xiaofei Li, Yutong Ban, Laurent Girin +2
This paper presents an online multiple-speaker localization and tracking method, as the INRIA-Perception contribution to the LOCATA Challenge 2018. First, the recursive least-squar…
Tracking Multiple Audio Sources with the von Mises Distribution and Variational EM
Yutong Ban, Xavier Alameda-PIneda, Christine Evers +1
In this paper we address the problem of simultaneously tracking several moving audio sources, namely the problem of estimating source trajectories from a sequence of observed featu…
Variational Bayesian Inference for Audio-Visual Tracking of Multiple Speakers
Yutong Ban, Xavier Alameda-Pineda, Laurent Girin +1
In this paper we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature of these two mo…
Online Localization and Tracking of Multiple Moving Speakers in Reverberant Environments
Xiaofei Li, Yutong Ban, Laurent Girin +2
We address the problem of online localization and tracking of multiple moving speakers in reverberant environments. The paper has the following contributions. We use the direct-pat…
A Deep Network for Arousal-Valence Emotion Prediction with Acoustic-Visual Cues
Songyou Peng, Le Zhang, Yutong Ban +2
In this paper, we comprehensively describe the methodology of our submissions to the One-Minute Gradual-Emotion Behavior Challenge 2018.