4 papers
Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG
Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6
Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, whil…
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Hyunsoo Kim, Jungmyung Wi, Soobin Um +2
Prior studies have demonstrated that diffusion classifiers achieve robust zero-shot classification performance. However, their effectiveness is strongly tied to the pretraining dat…
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Jungmyung Wi, Hyunsoo Kim, Donghyun Kim
Text-to-image models produce images that align well with natural language prompts, but compositional generation has long been a central challenge. Models often struggle to satisfy…
Training-Free Label Space Alignment for Universal Domain Adaptation
Dujin Lee, Sojung An, Jungmyung Wi +2
Universal domain adaptation (UniDA) transfers knowledge from a labeled source domain to an unlabeled target domain, where label spaces may differ and the target domain may contain…