617 citations · 2.2k across the 129 of their papers we have counts for
7 papers · 2 filters
HyperCLIP: Adapting Vision-Language models with Hypernetworks
Victor Akinwande, Mohammad Sadegh Norouzzadeh, Devin Willmott +3
Self-supervised vision-language models trained with contrastive objectives form the basis of current state-of-the-art methods in AI vision tasks. The success of these models is a d…
Diffusing Differentiable Representations
Yash Savani, Marc Finzi, J. Zico Kolter
We introduce a novel, training-free method for sampling differentiable representations (diffreps) using pretrained diffusion models. Rather than merely mode-seeking, our method ach…
Blind Inverse Problem Solving Made Easy by Text-to-Image Latent Diffusion
Michail Dontas, Yutong He, Naoki Murata +3
This paper considers blind inverse image restoration, the task of predicting a target image from a degraded source when the degradation (i.e. the forward operator) is unknown. Exis…
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
Kevin Y. Li, Sachin Goyal, Joao D. Semedo +1
Vision Language Models (VLMs) have demonstrated strong capabilities across various visual understanding and reasoning tasks, driven by incorporating image representations into the…
One-Step Diffusion Distillation through Score Implicit Matching
Weijian Luo, Zemin Huang, Zhengyang Geng +2
Despite their strong performances on many generative tasks, diffusion models require a large number of sampling steps in order to generate realistic samples. This has motivated the…
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
Joshua Nathaniel Williams, Avi Schwarzschild, Yutong He +1
Recovering natural language prompts for image generation models, solely based on the generated images is a difficult discrete optimization problem. In this work, we present the fir…