activity
20162023
most citedClova Baseline System for the VoxCeleb Speaker Recognition Challenge 2020

97 citations · 328 across the 22 of their papers we have counts for

collaborators
Showing 2018 · cs.CVShow all

5 papers · 2 filters

cs.CV2018

Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, whe…

cs.CV2018

LRS3-TED: a large-scale dataset for visual speech recognition

Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the c…

cs.CV2018

Deep Audio-Visual Speech Recognition

Triantafyllos Afouras, Joon Son Chung, Andrew Senior +2

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a lim…

cs.CV2018

Deep Lip Reading: a comparison of models and an online application

Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

The goal of this paper is to develop state-of-the-art models for lip reading -- visual speech recognition. We develop three architectures and compare their accuracy and training ti…

cs.CV2018

The Conversation: Deep Audio-Visual Speech Enhancement

Triantafyllos Afouras, Joon Son Chung, Andrew Zisserman

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known sp…