11 papers
Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution
Nripsuta Ani Saxena, Stelios Triantafyllou, Goran Radanović
With the growing adoption of artificial intelligence in high-stakes decision-making, identifying the causes of outcomes--particularly failures--and determining who is responsible h…
CONTRA: Red-Teaming Configurations of Personalizable Agents
Jonathan Nöther, Adish Singla, Goran Radanovic
Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization of…
Distributionally Robust Reinforcement Learning with Human Feedback
Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-rob…
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan Nöther, Adish Singla, Goran Radanovic
LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fail or behave adversarially. In th…
Sparse Offline Reinforcement Learning with Corruption Robustness
Nam Phuong Tran, Andi Nika, Goran Radanovic +2
We investigate robustness to strong data corruption in offline sparse reinforcement learning (RL). In our setting, an adversary may arbitrarily perturb a fraction of the collected…
Counterfactual Effect Decomposition in Multi-Agent Sequential Decision Making
Stelios Triantafyllou, Aleksa Sukovic, Yasaman Zolfimoselo +1
We address the challenge of explaining counterfactual outcomes in multi-agent Markov decision processes. In particular, we aim to explain the total counterfactual effect of an agen…