1 paper
Min Cai, Yuchen Zhang, Shichang Zhang +5
We propose SelfControl, an inference-time model control method utilizing gradients to control the behavior of large language models (LLMs) without explicit human annotations. Given…