3 papers
cs.AI2026
FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning
Hyungyu Choi, Young Kyun Jang, Chanho Eom
Vision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-t…
cs.CV2026
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
Gyuwon Han, Young Kyun Jang, Chanho Eom
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing…
cs.CV2025
GOAL: Global-local Object Alignment Learning
Hyungyu Choi, Young Kyun Jang, Chanho Eom
Vision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of thei…