2 papers
stat.ML2026
GRANITE: A Generalized Regional Framework for Identifying Agreement in Feature-Based Explanations
Julia Herbinger, Gabriel Laberge, Maximilian Muschalik +3
Feature-based explanation methods aim to quantify how features influence the model's behavior, either locally or globally, but different methods often disagree, producing conflicti…
cs.AI2025
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
Mina Taraghi, Yann Pequignot, Amin Nikanjam +2
Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work shows that even fine-tuning on benign dat…