4 citations · 7 across the 5 of their papers we have counts for
1 paper · 1 filter
Atharvan Dogra, Krishna Pillutla, Ameet Deshpande +5
We explore the ability of large language models (LLMs) to engage in subtle deception through strategically phrasing and intentionally manipulating information. This harmful behavio…