1 paper
Narmeen Oozeer, Sinem Erisken, Alice Rigg
Efforts to interpret reinforcement learning (RL) models often rely on high-level techniques such as attribution or probing, which provide only correlational insights and coarse cau…