3 papers
cs.CV2026
Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG
Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6
Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, whil…
cs.AI2026
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Jungmyung Wi, Hyunsoo Kim, Donghyun Kim
Text-to-image models produce images that align well with natural language prompts, but compositional generation has long been a central challenge. Models often struggle to satisfy…
cs.CV2024
CREPE: Coordinate-Aware End-to-End Document Parser
Yamato Okamoto, Youngmin Baek, Geewook Kim +5
In this study, we formulate an OCR-free sequence generation model for visual document understanding (VDU). Our model not only parses text from document images but also extracts the…