3 papers
cs.SD2026
Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
Ilpo Viertola, Giulio Cengarle, Gouthaman KV +2
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the…
cs.CV2025
Video Object Segmentation-Aware Audio Generation
Ilpo Viertola, Vladimir Iashin, Esa Rahtu
Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on…
cs.CV2024
Temporally Aligned Audio for Video with Autoregression
Ilpo Viertola, Vladimir Iashin, Esa Rahtu
We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extra…