activity
20202026
most citedTaming Visually Guided Sound Generation

19 citations · 30 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2025

Video Object Segmentation-Aware Audio Generation

Ilpo Viertola, Vladimir Iashin, Esa Rahtu

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on…

cs.CV2025

Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder

Vladimir Iashin, Horace Lee, Dan Schofield +1

Camera traps are revolutionising wildlife monitoring by capturing vast amounts of visual data; however, the manual identification of individual animals remains a significant bottle…

cs.CV2024

Temporally Aligned Audio for Video with Autoregression

Ilpo Viertola, Vladimir Iashin, Esa Rahtu

We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extra…

cs.CV2024★ 1 cited

Synchformer: Efficient Synchronization from Sparse Cues

Vladimir Iashin, Weidi Xie, Esa Rahtu +1

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a…

cs.CV2022★ 2 cited

Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Vladimir Iashin, Weidi Xie, Esa Rahtu +1

The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spati…

cs.CV2021★ 19 cited

Taming Visually Guided Sound Generation

Vladimir Iashin, Esa Rahtu

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art…