26 citations · 89 across the 40 of their papers we have counts for
Showing 2026 · cs.CVShow all
2 papers · 2 filters
cs.CV2026
Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai +2
RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of both text and image generation. However, m…
cs.CV2026
Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment
Bac Nguyen, Yuhta Takida, Naoki Murata +4
Slot Attention (SA) with pretrained diffusion models has recently shown promise for object-centric learning (OCL), but suffers from slot entanglement and weak alignment between obj…