activity
20222024
most citedLearning in Audio-visual Context: A Review, Analysis, and New Perspective

32 citations · 102 across the 23 of their papers we have counts for

collaborators

23 papers

cs.HC20249 cited

MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos

Zheng Ning, Zheng Zhang, Jerrick Ban +4

Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and ski…

cs.CV2024

OSCaR: Object State Captioning and State Change Representation

Nguyen Nguyen, Jing Bi, Ali Vosoughi +3

The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human…

cs.MM2024

Robust Active Speaker Detection in Noisy Environments

Siva Sai Nagender Vasireddy, Chenxu Zhang, Xiaohu Guo +1

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leve…

cs.SD20242 cited

Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, t…

cs.CV20242 cited

Efficiently Leveraging Linguistic Priors for Scene Text Spotting

Nguyen Nguyen, Yapeng Tian, Chenliang Xu

Incorporating linguistic knowledge can improve scene text recognition, but it is questionable whether the same holds for scene text spotting, which typically involves text detectio…

cs.HC2024

SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers

Zheng Ning, Brianna L. Wimer, Kaiwen Jiang +5

Blind or Low-Vision (BLV) users often rely on audio descriptions (AD) to access video content. However, conventional static ADs can leave out detailed information in videos, impose…