7 papers
Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning
Rogerio Guimaraes, Pietro Perona
Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models.…
Representational Difference Explanations
Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona
We propose a method for discovering and visualizing the differences between two learned representations, enabling more direct and interpretable model comparisons. We validate our m…
SAVeD: Learning to Denoise Low-SNR Video for Improved Downstream Performance
Suzanne Stathatos, Michael Hobley, Pietro Perona +1
Low signal-to-noise ratio videos -- such as those from underwater sonar, ultrasound, and microscopy -- pose significant challenges for computer vision models, particularly when pai…
Diffusion-Based Action Recognition Generalizes to Untrained Domains
Rogerio Guimaraes, Frank Xiao, Pietro Perona +1
Humans can recognize the same actions despite large context and viewpoint variations, such as differences between species (walking in spiders vs. horses), viewpoints (egocentric vs…
Social Perception of Faces in a Vision-Language Model
Carina I. Hausladen, Manuel Knott, Colin F. Camerer +1
We explore social perception of human faces in CLIP, a widely used open-source vision-language model. To this end, we compare the similarity in CLIP embeddings between different te…
Representational Similarity via Interpretable Visual Concepts
Neehar Kondapaneni, Oisin Mac Aodha, Pietro Perona
How do two deep neural networks differ in how they arrive at a decision? Measuring the similarity of deep networks has been a long-standing open question. Most existing methods pro…