5 papers
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
Trang Nguyen, Sidong Zhang, Shiv Shankar +4
Understanding and forecasting audience reactions to video content are crucial for improving content creation, recommendation systems, and media analysis. To enable audience reactio…
Resolution-Aware Retrieval Augmented Zero-Shot Forecasting
Iman Deznabi, Peeyush Kumar, Madalina Fiterau
Zero-shot forecasting aims to predict outcomes for previously unseen conditions without direct historical data, posing a significant challenge for traditional forecasting methods.…
Challenges in Understanding Modality Conflict in Vision-Language Models
Trang Nguyen, Jackson Michaels, Madalina Fiterau +1
This paper highlights the challenge of decomposing conflict detection from conflict resolution in Vision-Language Models (VLMs) and presents potential approaches, including using a…
Audio-Visual Speech Separation via Bottleneck Iterative Network
Sidong Zhang, Shiv Shankar, Trang Nguyen +2
Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to ob…
Leveraging Foundation Language Models (FLMs) for Automated Cohort Extraction from Large EHR Databases
Purity Mugambi, Alexandra Meliou, Madalina Fiterau
A crucial step in cohort studies is to extract the required cohort from one or more study datasets. This step is time-consuming, especially when a researcher is presented with a da…