2 papers
cs.CV2025
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
Yehna Kim, Young-Eun Kim, Seong-Whan Lee
Vision-Language Models (VLMs) have demonstrated impressive capabilities in zero-shot action recognition by learning to associate video embeddings with class embeddings. However, a…
cs.CV2025
FAR-Net: Multi-Stage Fusion Network with Enhanced Semantic Alignment and Adaptive Reconciliation for Composed Image Retrieval
Jeong-Woo Park, Young-Eun Kim, Seong-Whan Lee
Composed image retrieval (CIR) is a vision language task that retrieves a target image using a reference image and modification text, enabling intuitive specification of desired ch…