3 papers
cs.CV2025
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
Tiantian Geng, Teng Wang, Jinming Duan +4
Video event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods tend to over-spe…
cs.CV2025
Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
Xingyu Miao, Haoran Duan, Yang Bai +5
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on mult…
cs.CV2024
Relation-Aware Meta-Learning for Zero-shot Sketch-Based Image Retrieval
Yang Liu, Jiale Du, Xinbo Gao +1
Sketch-based image retrieval (SBIR) relies on free-hand sketches to retrieve natural photos within the same class. However, its practical application is limited by its inability to…