3 papers
cs.CV2025
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
Xinhao Xiang, Kuan-Chuan Peng, Suhas Lohit +2
3D object detection plays a crucial role in autonomous systems, yet existing methods are limited by closed-set assumptions and struggle to recognize novel objects and their attribu…
cs.CV2025
Programmatic Video Prediction Using Large Language Models
Hao Tang, Kevin Ellis, Suhas Lohit +2
The task of estimating the world model describing the dynamics of a real world process assumes immense importance for anticipating and preparing for future outcomes. For applicatio…
cs.CV2025
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang +3
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a…