collaborators

6 papers

cs.LG2026

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

Ram Rachum, Yotam Amitai, Bálint Gyevnár +2

This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like…

cs.CR2026

Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud

Bálint Gyevnár, Atoosa Kasirzadeh, Nihar B. Shah

Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required the resources of a company: deep pockets,…

cs.CY2026

Bridging the Gap in the Responsible AI Divides

Bálint Gyevnár, Atoosa Kasirzadeh

Tensions between AI Safety (AIS) and AI Ethics (AIE) have increasingly surfaced in AI governance and public debates about AI, leading to what we term the "responsible AI divides".…

cs.AI2025

Integrating Counterfactual Simulations with Language Models for Explaining Multi-Agent Behaviour

Bálint Gyevnár, Christopher G. Lucas, Stefano V. Albrecht +1

Autonomous multi-agent systems (MAS) are useful for automating complex tasks but raise trust concerns due to risks such as miscoordination or goal misalignment. Explainability is v…

cs.HC2025

People Attribute Purpose to Autonomous Vehicles When Explaining Their Behavior: Insights from Cognitive Science for Explainable AI

Balint Gyevnar, Stephanie Droop, Tadeg Quillien +4

It is often argued that effective human-centered explainable artificial intelligence (XAI) should resemble human reasoning. However, empirical investigations of how concepts from c…

cs.AI2025

Objective Metrics for Human-Subjects Evaluation in Explainable Reinforcement Learning

Balint Gyevnar, Mark Towers

Explanation is a fundamentally human process. Understanding the goal and audience of the explanation is vital, yet existing work on explainable reinforcement learning (XRL) routine…