Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When can we trust untrusted monitoring? A safety case sketch across collusion strategies
Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov +4
AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitori…
cs.AI2024
'Explaining RL Decisions with Trajectories': A Reproducibility Study
Karim Abdel Sadek, Matteo Nulli, Joan Velja +1
This work investigates the reproducibility of the paper 'Explaining RL decisions with trajectories'. The original paper introduces a novel approach in explainable reinforcement lea…