From the 1 of 10 linked papers with an AI index.
10 papers
Leveraging Prior Knowledge of Diffusion Model for Person Search
Giyeol Kim, Sooyoung Yang, Jihyong Oh +2
The paper presents DiffPS, a person search framework that incorporates a pre-trained diffusion model to improve both person detection and re-identification, using three specialized…
FLAIR: Frequency- and Locality-Aware Implicit Neural Representations
Sukhun Ko, Seokhyun Youn, Dahyeon Kye +3
Implicit Neural Representations (INRs) leverage neural networks to map coordinates to corresponding signals, enabling continuous and compact representations. This paradigm has driv…
SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset
Changhyun Roh, Yonghyun Jeong, Jonghyun Lee +2
Synthesizing a target concept from a single reference image is challenging in diffusion-based personalized text-to-image generation, particularly for sticker personalization where…
AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
Dahyeon Kye, Changhyun Roh, Sukhun Ko +2
Video Frame Interpolation (VFI) is a core low-level vision task that synthesizes intermediate frames between existing ones while ensuring spatial and temporal coherence. Over the p…
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
Gyuwon Han, Young Kyun Jang, Chanho Eom
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing…
Visual Representation Alignment for Multimodal Large Language Models
Heeji Yoon, Jaewoo Jung, Junwan Kim +10
Multimodal large language models (MLLMs) trained with visual instruction tuning have achieved strong performance across diverse tasks, yet they remain limited in vision-centric tas…