1 citations · 1 across the 3 of their papers we have counts for
6 papers
Inplace knowledge distillation with teacher assistant for improved training of flexible deep neural networks
Alexey Ozerov, Ngoc Duong
Deep neural networks (DNNs) have achieved great success in various machine learning tasks. However, most existing powerful DNN models are computationally expensive and memory deman…
On the hidden treasure of dialog in video question answering
Deniz Engin, François Schnitzler, Ngoc Q. K. Duong +1
High-level understanding of stories in video such as movies and TV shows from raw data is extremely challenging. Modern video question answering (VideoQA) systems often use additio…
Discriminate natural versus loudspeaker emitted speech
Thanh-Ha Le, Philippe Gilberton, Ngoc Q. K. Duong
In this work, we address a novel, but potentially emerging, problem of discriminating the natural human voices and those played back by any kind of audio devices in the context of…
Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision
Sanjeel Parekh, Alexey Ozerov, Slim Essid +3
We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object…
Weakly Supervised Representation Learning for Unsynchronized Audio-Visual Events
Sanjeel Parekh, Slim Essid, Alexey Ozerov +3
Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel…
A Review of Audio Features and Statistical Models Exploited for Voice Pattern Design
Ngoc Q. K. Duong, Hien-Thanh Duong
Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major…