activity
20242026
collaborators

11 papers

cs.SD2026

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6

Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, s…

cs.LG2026

GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

Naoki Murata, Yuhta Takida, Chieh-Hsin Lai +4

Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods score individual examples, practitioners o…

cs.LG2026

A Unified View of Score-Based and Drifting Models

Chieh-Hsin Lai, Bac Nguyen, Naoki Murata +5

Drifting models train one-step generators by optimizing a kernel-induced mean-shift discrepancy between the data and model distributions, with Laplace kernels used by default in pr…

cs.CV2026

PAVAS: Physics-Aware Video-to-Audio Synthesis

Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka +2

Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing vis…

cs.CV2026

Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment

Bac Nguyen, Yuhta Takida, Naoki Murata +4

Slot Attention (SA) with pretrained diffusion models has recently shown promise for object-centric learning (OCL), but suffers from slot entanglement and weak alignment between obj…

cs.LG2025

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida +3

In this study, we propose an enhancement to the similarity computation mechanism in multi-modal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demo…