617 citations · 2.2k across the 129 of their papers we have counts for
5 papers · 2 filters
One-Step Diffusion Distillation via Deep Equilibrium Models
Zhengyang Geng, Ashwini Pokle, J. Zico Kolter
Diffusion models excel at producing high-quality samples but naively require hundreds of iterations, prompting multiple attempts to distill the generation process into a faster net…
Text Descriptions are Compressive and Invariant Representations for Visual Learning
Zhili Feng, Anna Bair, J. Zico Kolter
Modern image classification is based upon directly predicting classes via large discriminative networks, which do not directly contain information about the intuitive visual featur…
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
Pratyush Maini, Sachin Goyal, Zachary C. Lipton +2
Large web-sourced multimodal datasets have powered a slew of new methods for learning general-purpose visual representations, advancing the state of the art in computer vision and…
Localized Text-to-Image Generation for Free via Cross Attention Control
Yutong He, Ruslan Salakhutdinov, J. Zico Kolter
Despite the tremendous success in text-to-image generative models, localized text-to-image generation (that is, generating objects or features at specific locations in an image whi…
Mimetic Initialization of Self-Attention Layers
Asher Trockman, J. Zico Kolter
It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point. We explore the weights of such pre-…