2 papers
cs.CV2026
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
Yerim So, Jiyeong Kim, Jiwon Yoon +1
Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignment, a process that obscures l…
cs.CV2025
Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection
Soyul Lee, Seungmin Baek, Dongbo Min
Monocular 3D object detection is a cost-effective solution for applications like autonomous driving and robotics, but remains fundamentally ill-posed due to inherently ambiguous de…