3 papers
cs.CV2026
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
Trang Nguyen, Sidong Zhang, Shiv Shankar +4
Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reactio…
cs.LG2025
Challenges in Understanding Modality Conflict in Vision-Language Models
Trang Nguyen, Jackson Michaels, Madalina Fiterau +1
This paper highlights the challenge of decomposing conflict detection from conflict resolution in Vision-Language Models (VLMs) and presents potential approaches, including using a…
cs.SD2025
Audio-Visual Speech Separation via Bottleneck Iterative Network
Sidong Zhang, Shiv Shankar, Trang Nguyen +2
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to ob…