4 papers
GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling
Shivanshu Shekhar, Uttaran Bhattacharya, Raghavendra Addanki +3
Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward modeling, but these models struggle…
VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion
Shivanshu Shekhar, Sagnik Mukherjee, Jia Yi Zhang +1
Sequential Monte Carlo (SMC) samplers for reward-guided diffusion models often suffer from rapid lineage collapse: a few high-reward particles dominate the population within a hand…
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
Shivanshu Shekhar, Shreyas Singh, Tong Zhang
Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to…
ROCM: RLHF on consistency models
Shivanshu Shekhar, Tong Zhang
Diffusion models have revolutionized generative modeling in continuous domains like image, audio, and video synthesis. However, their iterative sampling process leads to slow gener…