13 papers
DRM: Diffusion-based Reward Model With Step-wise Guidance
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alig…
Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition
Chencheng Tang, Xuanyu Xue, Fangyikang Wang +2
Diffusion models have achieved impressive empirical success in generative tasks, and their convergence theory is now relatively well understood. Motivated by privacy and scalabilit…
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
Shibei Meng, Binxin Yang, Yuan Liu +4
Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such pointwise supervision often mi…
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
Zeyue Tian, Binxin Yang, Zhaoyang Liu +8
Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized…
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
Tianlin Pan, Jiayi Dai, Chenpu Yuan +7
Recent video editing models have achieved impressive results, but most still require large-scale paired datasets. Collecting such naturally aligned pairs at scale remains highly ch…
Improving Reconstruction of Representation Autoencoder
Siyu Liu, Chujie Qin, Hubery Yin +6
Recent work leverages Vision Foundation Models as image encoders to boost the generative performance of latent diffusion models (LDMs), as their semantic feature distributions are…