3 papers
cs.LG2026
Towards Understanding Steering Strength
Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations. Namely, identify a well-chosen direction dependi…
cs.LG2025
Feature Attribution from First Principles
Magamed Taimeskhanov, Damien Garreau
Feature attribution methods are a popular approach to explain the behavior of machine learning models. They assign importance scores to each input feature, quantifying their influe…
cs.CV2024
CAM-Based Methods Can See through Walls
Magamed Taimeskhanov, Ronan Sicre, Damien Garreau
CAM-based methods are widely-used post-hoc interpretability method that produce a saliency map to explain the decision of an image classification model. The saliency map highlights…