32 papers
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Yuya Kobayashi, Masato Ishii, Yuhta Takida +2
Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies b…
From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
Luca Ambrogioni, Giulio Franzese, Alberto Foresti +7
How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or…
TILDE: TILt-based Distributional Erasure for Concept Unlearning
Naveen George, Naoki Murata, Yuhta Takida +2
Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes, trademark constraints, and saf…
Locality-Aware Continual Unlearning for Diffusion Models
Naveen George, Naoki Murata, Yuhta Takida +2
Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations arise over time. Existing unlearning met…
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, s…
Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model
Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai +2
RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of both text and image generation. However, m…