Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Youngkil Song, Yoonjae Baek, Dongwon Kim +3
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling high-level reasoning with temporal groun…
cs.CV2024
Improving Text-based Person Search via Part-level Cross-modal Correspondence
Jicheol Park, Boseung Jeong, Dongwon Kim +1
Text-based person search is the task of finding person images that are the most relevant to the natural language text description given as query. The main challenge of this task is…
cs.CV2024
Bootstrapping Top-down Information for Self-modulating Slot Attention
Dongwon Kim, Seoyeon Kim, Suha Kwak
Object-centric learning (OCL) aims to learn representations of individual objects within visual scenes without manual supervision, facilitating efficient and effective visual reaso…