1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
Jiaxin Wen, Vivek Hebbar, Caleb Larson +9
As large language models (LLMs) become increasingly capable, it is prudent to assess whether safety measures remain effective even if LLMs intentionally try to bypass them. Previou…
cs.CL2024
Spontaneous Reward Hacking in Iterative Self-Refinement
Jane Pan, He He, Samuel R. Bowman +1
Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human u…