2 papers
cs.CV2025
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang +3
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a…
cs.CV2024
Understanding and Improving Training-Free AI-Generated Image Detections with Vision Foundation Models
Chung-Ting Tsai, Ching-Yun Ko, I-Hsin Chung +2
The rapid advancement of generative models has introduced serious risks, including deepfake techniques for facial synthesis and editing. Traditional approaches rely on training cla…