activity
20152021
most citedA Review of Audio Features and Statistical Models Exploited for Voice Pattern Design

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

eess.SP2021

Inplace knowledge distillation with teacher assistant for improved training of flexible deep neural networks

Alexey Ozerov, Ngoc Duong

Deep neural networks (DNNs) have achieved great success in various machine learning tasks. However, most existing powerful DNN models are computationally expensive and memory deman…

cs.CV2021

On the hidden treasure of dialog in video question answering

Deniz Engin, François Schnitzler, Ngoc Q. K. Duong +1

High-level understanding of stories in video such as movies and TV shows from raw data is extremely challenging. Modern video question answering (VideoQA) systems often use additio…

cs.SD2019

Discriminate natural versus loudspeaker emitted speech

Thanh-Ha Le, Philippe Gilberton, Ngoc Q. K. Duong

In this work, we address a novel, but potentially emerging, problem of discriminating the natural human voices and those played back by any kind of audio devices in the context of…

cs.CV2018

Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision

Sanjeel Parekh, Alexey Ozerov, Slim Essid +3

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object…

cs.CV2018

Weakly Supervised Representation Learning for Unsynchronized Audio-Visual Events

Sanjeel Parekh, Slim Essid, Alexey Ozerov +3

Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel…

cs.SD20151 cited

A Review of Audio Features and Statistical Models Exploited for Voice Pattern Design

Ngoc Q. K. Duong, Hien-Thanh Duong

Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major…