most citedTaming Visually Guided Sound Generation

19 citations · 21 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2022

Supervised Fine-tuning Evaluation for Long-term Visual Place Recognition

Farid Alijani, Esa Rahtu

In this paper, we present a comprehensive study on the utility of deep convolutional neural networks with two state-of-the-art pooling layers which are placed after convolutional l…

cs.CV20222 cited

Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Vladimir Iashin, Weidi Xie, Esa Rahtu +1

The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spati…

cs.CV202119 cited

Taming Visually Guided Sound Generation

Vladimir Iashin, Esa Rahtu

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art…

cs.CV2021

V-SlowFast Network for Efficient Visual Sound Separation

Lingyu Zhu, Esa Rahtu

The objective of this paper is to perform visual sound separation: i) we study visual sound separation on spectrograms of different temporal resolutions; ii) we propose a new light…

cs.CV2021

Lightweight Monocular Depth with a Novel Neural Architecture Search Method

Lam Huynh, Phong Nguyen, Jiri Matas +2

This paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture sear…

cs.CV2021

Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian Loss

Lam Huynh, Matteo Pedone, Phong Nguyen +3

Deep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy a…