collaborators

6 papers

cs.LG2026

RelShap: Relationally Consistent Shapley Explanations

Seungeun Lee, Joao Fonseca, Julia Stoyanovich

Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature attributio…

cs.LG2026

ExplainerPFN: Towards tabular foundation models for model-free zero-shot feature importance estimations

Joao Fonseca, Julia Stoyanovich

Computing the importance of features in supervised classification tasks is critical for model interpretability. Shapley values are a widely used approach for explaining model predi…

cs.AI2025

ShaRP: Explaining Rankings and Preferences with Shapley Values

Venetia Pliatsika, Joao Fonseca, Kateryna Akhynko +2

Algorithmic decisions in critical domains such as hiring, college admissions, and lending are often based on rankings. Given the impact of these decisions on individuals, organizat…

cs.LG2025

SHAP-based Explanations are Sensitive to Feature Representation

Hyunseung Hwang, Andrew Bell, Joao Fonseca +3

Local feature-based explanations are a key component of the XAI toolkit. These explanations compute feature importance values relative to an ``interpretable'' feature representatio…

cs.CL2025

Output Scouting: Auditing Large Language Models for Catastrophic Responses

Andrew Bell, Joao Fonseca

Recent high profile incidents in which the use of Large Language Models (LLMs) resulted in significant harm to individuals have brought about a growing interest in AI safety. One r…

cs.CL2025

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs

Joao Fonseca, Andrew Bell, Julia Stoyanovich

Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model. Jailbreaks have been ex…