4 papers
PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
Hee Suk Yoon, Eunseop Yoon, Ji Woo Hong +6
Reinforcement Learning with Verifiable Rewards (RLVR) traditionally relies on a sparse, outcome-based signal. Recent work shows that providing a fine-grained, model-intrinsic signa…
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo +5
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image to…
E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization
Trung X. Pham, Zhang Kang, Ji Woo Hong +2
We propose E-MD3C (fficient asked iffusion Transformer with Disentangled onditions and ompact $\underline…
DriVLM: Domain Adaptation of Vision-Language Models in Autonomous Driving
Xuran Zheng, Chang D. Yoo
In recent years, large language models have had a very impressive performance, which largely contributed to the development and application of artificial intelligence, and the para…