3 papers
cs.CV2026
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
Trang Nguyen, Sidong Zhang, Shiv Shankar +4
Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reactio…
cs.SD2025
Audio-Visual Speech Separation via Bottleneck Iterative Network
Sidong Zhang, Shiv Shankar, Trang Nguyen +2
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to ob…
cs.LG2025
Learning Straight Flows by Learning Curved Interpolants
Shiv Shankar, Tomas Geffner
Flow matching models typically use linear interpolants to define the forward/noise addition process. This, together with the independent coupling between noise and target distribut…