8 papers
Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model
Anirud Aggarwal, Abhinav Shrivastava, Matthew Gwilliam
Diffusion-based image generation models excel at producing high-quality synthetic content, but suffer from slow and computationally expensive inference. Prior work has attempted to…
Towards Understanding Best Practices for Quantization of Vision-Language Models
Gautom Das, Vincent La, Ethan Lau +2
Large language models (LLMs) deliver impressive results for a variety of tasks, but state-of-the-art systems require fast GPUs with large amounts of memory. To reduce both the memo…
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
Eric Zhu, Mara Levy, Matthew Gwilliam +1
Training a policy that can generalize to unknown objects is a long standing challenge within the field of robotics. The performance of a policy often drops significantly in situati…
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
Vatsal Agarwal, Matthew Gwilliam, Gefen Kohavi +3
Recent advances in multimodal large language models (MLLMs) have enabled image-based question-answering capabilities. However, a key limitation is the use of CLIP as the visual enc…
How to Design and Train Your Implicit Neural Representation for Video Compression
Matthew Gwilliam, Roy Zhang, Namitha Padmanabhan +2
Implicit neural representation (INR) methods for video compression have recently achieved visual quality and compression ratios that are competitive with traditional pipelines. How…
Accelerate High-Quality Diffusion Models with Inner Loop Feedback
Matthew Gwilliam, Han Cai, Di Wu +2
We propose Inner Loop Feedback (ILF), a novel approach to accelerate diffusion models' inference. ILF trains a lightweight module to predict future features in the denoising proces…