18 citations · 18 across the 4 of their papers we have counts for
1 paper · 1 filter
Mariana Lins Costa
The prevailing technical literature in AI Safety interprets scheming and sandbagging behaviors in large language models (LLMs) as indicators of deceptive agency or hidden objective…