1 paper
Sagi Polaczek, Noa Kraicer, Gal Metzer +4
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leaka…