Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang +3
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a…
cs.CV2025
Improving Open-World Object Localization by Discovering Background
Ashish Singh, Michael J. Jones, Kuan-Chuan Peng +3
Our work addresses the problem of learning to localize objects in an open-world setting, i.e., given the bounding box information of a limited number of object classes during train…
cs.CV2023
Robust Frame-to-Frame Camera Rotation Estimation in Crowded Scenes
Fabien Delattre, David Dirnfeld, Phat Nguyen +4
We present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no pre…