activity
20172026
most citedWatching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows

2 citations · 3 across the 7 of their papers we have counts for

collaborators

11 papers

cs.CV2026

Spatially-Grounded Flow Matching: Structured Source Distributions for Image Generation

Arman Zarei, Mahdi M. Kalayeh

Current flow matching models learn to transport the source i.i.d. Gaussian noise into the target distribution of natural images, yet this source distribution carries no notion of s…

cs.CV2026

CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

Anku Rani, Wei Dai, Shravan Nayak +3

As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metri…

cs.CV2026

Envisioning the Future, One Step at a Time

Stefan Andreas Baumann, Jannik Wiese, Tommaso Martorella +2

Accurately anticipating how complex, diverse scenes will evolve requires models that represent uncertainty, simulate along extended interaction chains, and efficiently explore many…

cs.SD2025

Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation

Darius Petermann, Mahdi M. Kalayeh

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild vid…

cs.SD2023★ 1 cited

Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning

Nikhil Singh, Chih-Wei Wu, Iroro Orife +1

Audiovisual representation learning typically relies on the correspondence between sight and sound. However, there are often multiple audio tracks that can correspond with a visual…

cs.CV2022

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu +2

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an in…