1 paper · 1 filter
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai +2
RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of both text and image generation. However, m…