1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
SHAP-based Explanations are Sensitive to Feature Representation
Hyunseung Hwang, Andrew Bell, Joao Fonseca +3
Local feature-based explanations are a key component of the XAI toolkit. These explanations compute feature importance values relative to an ``interpretable'' feature representatio…
cs.CL2025
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Joao Fonseca, Andrew Bell, Julia Stoyanovich
Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model. Jailbreaks have been ex…
cs.CY2024
Making Transparency Advocates: An Educational Approach Towards Better Algorithmic Transparency in Practice
Andrew Bell, Julia Stoyanovich
Concerns about the risks and harms posed by artificial intelligence (AI) have resulted in significant study into algorithmic transparency, giving rise to a sub-field known as Expla…