4 papers · 1 filter
Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG
Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6
Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, whil…
Prompt Learning via Meta-Regularization
Jinyoung Park, Juyeon Ko, Hyunwoo J. Kim
Pre-trained vision-language models have shown impressive success on various computer vision tasks with their zero-shot generalizability. Recently, prompt learning approaches have b…
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis
Juyeon Ko, Inho Kong, Dogyun Park +1
Semantic image synthesis (SIS) is a task to generate realistic images corresponding to semantic maps (labels). However, in real-world applications, SIS often encounters noisy user…
Semantic-Aware Implicit Template Learning via Part Deformation Consistency
Sihyeon Kim, Minseok Joo, Jaewon Lee +3
Learning implicit templates as neural fields has recently shown impressive performance in unsupervised shape correspondence. Despite the success, we observe current approaches, whi…