3 papers
cs.LG2026
Decomposing multimodal embedding spaces with group-sparse autoencoders
Chiraag Kaushik, Davis Barch, Andrea Fanelli
The Linear Representation Hypothesis asserts that the embeddings learned by neural networks can be understood as linear combinations of features corresponding to high-level concept…
cs.SD2025
Audio-Visual Speech Separation via Bottleneck Iterative Network
Sidong Zhang, Shiv Shankar, Trang Nguyen +2
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to ob…
cs.MM2024
Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
Shanti Stewart, Gouthaman KV, Lie Lu +1
Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best musi…