activity
20222026
most citedConsistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

16 citations · 39 across the 43 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Spectral Prior for Reducing Exposure Bias in Diffusion Models

Yuya Kobayashi, Masato Ishii, Yuhta Takida +2

Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies b…

cs.CV2026

Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model

Yoonjeon Kim, Yuhta Takida, Chieh-Hsin Lai +2

RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of both text and image generation. However, m…

cs.CV2026

Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment

Bac Nguyen, Yuhta Takida, Naoki Murata +4

Slot Attention (SA) with pretrained diffusion models has recently shown promise for object-centric learning (OCL), but suffers from slot entanglement and weak alignment between obj…

cs.CV2025

PAVAS: Physics-Aware Video-to-Audio Synthesis

Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka +2

Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing vis…

cs.CV2025

Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution

Yonghyun Park, Chieh-Hsin Lai, Satoshi Hayakawa +7

While diffusion models excel at image generation, their growing adoption raises critical concerns about copyright issues and model transparency. Existing attribution methods identi…

cs.CV2025

Efficiency without Compromise: CLIP-aided Text-to-Image GANs with Increased Diversity

Yuya Kobayashi, Yuhta Takida, Takashi Shibuya +1

Recently, Generative Adversarial Networks (GANs) have been successfully scaled to billion-scale large text-to-image datasets. However, training such models entails a high training…