6 papers
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents
Ram Rachum, Yotam Amitai, Bálint Gyevnár +2
This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…
Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud
Bálint Gyevnár, Atoosa Kasirzadeh, Nihar B. Shah
Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets,…
Bridging the Gap in the Responsible AI Divides
Bálint Gyevnár, Atoosa Kasirzadeh
Tensions between AI Safety (AIS) and AI Ethics (AIE) have increasingly surfaced in AI governance and public debates about AI, leading to what we term the "responsible AI divides".…
Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour
Bálint Gyevnár, Christopher G. Lucas, Stefano V. Albrecht +1
Autonomous multi-agent systems (MAS) are useful for automating complex tasks but raise trust concerns due to risks such as miscoordination or goal misalignment. Explainability is v…
People Attribute Purpose to Autonomous Vehicles When Explaining Their Behavior: Insights from Cognitive Science for Explainable AI
Balint Gyevnar, Stephanie Droop, Tadeg Quillien +4
It is often argued that effective human-centered explainable artificial intelligence (XAI) should resemble human reasoning. However, empirical investigations of how concepts from c…
Objective Metrics for Human-Subjects Evaluation in Explainable Reinforcement Learning
Balint Gyevnar, Mark Towers
Explanation is a fundamentally human process. Understanding the goal and audience of the explanation is vital, yet existing work on explainable reinforcement learning (XRL) routine…