32 citations · 102 across the 23 of their papers we have counts for
23 papers
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
Zheng Ning, Zheng Zhang, Jerrick Ban +4
Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and ski…
OSCaR: Object State Captioning and State Change Representation
Nguyen Nguyen, Jing Bi, Ali Vosoughi +3
The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human…
Robust Active Speaker Detection in Noisy Environments
Siva Sai Nagender Vasireddy, Chenxu Zhang, Xiaohu Guo +1
This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leve…
Text-to-Audio Generation Synchronized with Videos
Shentong Mo, Jing Shi, Yapeng Tian
In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
Nguyen Nguyen, Yapeng Tian, Chenliang Xu
Incorporating linguistic knowledge can improve scene text recognition, but it is questionable whether the same holds for scene text spotting, which typically involves text detectio…
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
Zheng Ning, Brianna L. Wimer, Kaiwen Jiang +5
Blind or Low-Vision (BLV) users often rely on audio descriptions (AD) to access video content. However, conventional static ADs can leave out detailed information in videos, impose…