5 papers
Generalizing the Geometry of Model Merging Through Frechet Averages
Marvin F. da Silva, Mohammed Adnan, Felix Dangel +1
Model merging aims to combine multiple models into one without additional training. Naïve parameter-space averaging can be fragile under architectural symmetries, as their geometr…
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
Scott C. Lowe, Anthony Fuller, Sageev Oore +2
The landscape of self-supervised learning (SSL) is currently dominated by generative approaches (e.g. MAE) that reconstruct raw low-level data, and predictive approaches (e.g. I-JE…
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
Marvin F. da Silva, Felix Dangel, Sageev Oore
The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work repor…
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
Sri Harsha Dumpala, David Arps, Sageev Oore +2
Vision-language models (VLMs), serve as foundation models for multi-modal applications such as image captioning and text-to-image generation. Recent studies have highlighted limita…
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Sastry +3
Despite the significant influx of prompt-tuning techniques for generative vision-language models (VLMs), it remains unclear how sensitive these models are to lexical and semantic a…