4 papers
Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding
Hyomin Kim, Junghye Kim, Joanie Hayoun Chung +4
Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural weaknesses in existing reaso…
FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation
Yongjin Kim, Yoonjin Oh, Yerin Kim +5
With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, des…
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Donghwan Chi, Hyomin Kim, Yoonjin Oh +7
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been devel…
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
Yoonjin Oh, Yongjin Kim, Hyomin Kim +2
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image…