11 papers
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, s…
GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning
Naoki Murata, Yuhta Takida, Chieh-Hsin Lai +4
Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods score individual examples, practitioners o…
A Unified View of Score-Based and Drifting Models
Chieh-Hsin Lai, Bac Nguyen, Naoki Murata +5
Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with Laplace kernels used by default in pr…
PAVAS: Physics-Aware Video-to-Audio Synthesis
Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka +2
Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing vis…
Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment
Bac Nguyen, Yuhta Takida, Naoki Murata +4
Slot Attention (SA) with pretrained diffusion models has recently shown promise for object-centric learning (OCL), but suffers from slot entanglement and weak alignment between obj…
Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida +3
In this study, we propose an enhancement to the similarity computation mechanism in multi-modal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demo…