4 papers
Visual Fingerprints for LLM Generation Comparison
Amal Alnouri, Andreas Hinterreiter, Christina Humer +2
Large language model (LLM) outputs arise from complex interactions among prompts, system instructions, model parameters, and architecture. We refer to specific configurations of th…
CafGa: Customizing Feature Attributions to Explain Language Models
Alan Boyle, Furui Cheng, Vilém Zouhar +1
Feature attribution methods, such as SHAP and LIME, explain machine learning model predictions by quantifying the influence of each input component. When applying feature attributi…
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
Furui Cheng, Vilém Zouhar, Robin Shing Moon Chan +3
Understanding the behavior of large language models (LLMs) is crucial for ensuring their safe and reliable use. However, existing explainable AI (XAI) methods for LLMs primarily re…
DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
Danqing Shi, Furui Cheng, Tino Weinkauf +2
Human preferences are widely used to align large language models (LLMs) through methods such as reinforcement learning from human feedback (RLHF). However, the current user interfa…