21 papers
Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai +2
RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of both text and image generation. However, m…
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
Naoki Murata, Yuhta Takida, Chieh-Hsin Lai +4
Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods score individual examples, practitioners o…
The Principles of Diffusion Models
Chieh-Hsin Lai, Yang Song, Dongjun Kim +2
This book presents the core principles that have guided the development of diffusion models, tracing their origins and showing how diverse formulations arise from shared mathematic…
A Unified View of Score-Based and Drifting Models
Chieh-Hsin Lai, Bac Nguyen, Naoki Murata +5
Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with Laplace kernels used by default in pr…
Understanding and Accelerating the Training of Masked Diffusion Language Models
Chunsan Hong, Sanghyun Lee, Chieh-Hsin Lai +5
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs are known to learn substantially more sl…
Improving Classifier-Free Guidance in Masked Diffusion: Low-Dim Theoretical Insights with High-Dim Impact
Kevin Rojas, Ye He, Chieh-Hsin Lai +3
Classifier-Free Guidance (CFG) is a widely used technique for conditional generation and improving sample quality in continuous diffusion models, and its extensions to discrete dif…