evaluation metrics 1generalization across domains 1ideological bias 1language model finetuning 1model alignment 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Robert Graham, Edward Stevinson, Yariv Barsheshat
The paper shows that fine‑tuning large language models on small, factually defensible datasets can cause broad ideological shifts across unrelated topics, and introduces metrics to…
cs.CV2025
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
Sonia Joseph, Praneet Suresh, Lorenz Hufe +7
Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision…
cs.CV2025
Steering CLIP's vision transformer with sparse autoencoders
Sonia Joseph, Praneet Suresh, Ethan Goldfarb +6
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but whic…