Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1
Large-scale pre-trained vision-language models like CLIP demonstrate remarkable zero-shot performance across diverse tasks. However, fine-tuning these models to improve downstream…
cs.CV2025
On the rankability of visual embeddings
Ankit Sonthalia, Arnas Uselis, Seong Joon Oh
We study whether visual embedding models capture continuous, ordinal attributes along linear directions, which we term _rank axes_. We define a model as _rankable_ for an attribute…