62 citations · 78 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2022
Audio-Visual Fusion Layers for Event Type Aware Video Recognition
Arda Senocak, Junsik Kim, Tae-Hyun Oh +3
Human brain is continuously inundated with the multisensory information and their complex interactions coming from the outside world at any given moment. Such information is automa…
cs.CV2020★ 9 cited
Unified Multisensory Perception: Weakly-Supervised Audio-Visual Video Parsing
Yapeng Tian, Dingzeyu Li, Chenliang Xu
In this paper, we introduce a new problem, named audio-visual video parsing, which aims to parse a video into temporal event segments and label them as either audible, visible, or…
cs.CV2020
MakeItTalk: Speaker-Aware Talking-Head Animation
Yang Zhou, Xintong Han, Eli Shechtman +3
We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct…