1 paper
Gilad Gressel, Rahul Pankajakshan, Julia Diament +3
As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their behavior. This is difficult when models in…