1 paper
Zixuan Weng, Jinghuai Zhang, Kunlin Cai +3
Large language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations. Although inference-time steering offers a cost-effective way to adju…