Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Towards Understanding Steering Strength
Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction dependi…
cs.LG2025
Feature Attribution from First Principles
Magamed Taimeskhanov, Damien Garreau
Feature attribution methods are a popular approach to explain the behavior of machine learning models. They assign importance scores to each input feature, quantifying their influe…