6 papers
RelShap: Relationally Consistent Shapley Explanations
Seungeun Lee, Joao Fonseca, Julia Stoyanovich
Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature attributio…
ExplainerPFN: Towards tabular foundation models for model-free zero-shot feature importance estimations
Joao Fonseca, Julia Stoyanovich
Computing the importance of features in supervised classification tasks is critical for model interpretability. Shapley values are a widely used approach for explaining model predi…
ShaRP: Explaining Rankings and Preferences with Shapley Values
Venetia Pliatsika, Joao Fonseca, Kateryna Akhynko +2
Algorithmic decisions in critical domains such as hiring, college admissions, and lending are often based on rankings. Given the impact of these decisions on individuals, organizat…
SHAP-based Explanations are Sensitive to Feature Representation
Hyunseung Hwang, Andrew Bell, Joao Fonseca +3
Local feature-based explanations are a key component of the XAI toolkit. These explanations compute feature importance values relative to an ``interpretable'' feature representatio…
Output Scouting: Auditing Large Language Models for Catastrophic Responses
Andrew Bell, Joao Fonseca
Recent high profile incidents in which the use of Large Language Models (LLMs) resulted in significant harm to individuals have brought about a growing interest in AI safety. One r…
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Joao Fonseca, Andrew Bell, Julia Stoyanovich
Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model. Jailbreaks have been ex…