21 papers
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion mode…
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
Anirud Aggarwal, Abhinav Shrivastava, Matthew Gwilliam
Diffusion-based image generation models excel at producing high-quality synthetic content, but suffer from slow and computationally expensive inference. Prior work has attempted to…
TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
Namitha Padmanabhan, Matthew Gwilliam, Abhinav Shrivastava
Implicit Neural Representations (INRs) have recently demonstrated impressive performance for video compression. However, since a separate INR must be overfit for each video, scalin…
UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
Matthew Walmer, Saksham Suri, Anirud Aggarwal +1
The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act…
Towards Understanding Best Practices for Quantization of Vision-Language Models
Gautom Das, Vincent La, Ethan Lau +2
Large language models (LLMs) deliver impressive results for a variety of tasks, but state-of-the-art systems require fast GPUs with large amounts of memory. To reduce both the memo…
Characterizing Motion Encoding in Video Diffusion Timesteps
Vatsal Baherwani, Yixuan Ren, Abhinav Shrivastava
Text-to-video diffusion models synthesize temporal motion and spatial appearance through iterative denoising, yet how motion is encoded across timesteps remains poorly understood.…