4 papers
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Haoming Xu, Weihong Xu, Zongrui Li +6
Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what to ignore. We study this ch…
CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer
Yue Zhao, Yujia Gong, Ruigang Liang +4
The widespread deployment of large language models (LLMs) calls for post-hoc methods that can flexibly adapt models to evolving safety requirements. Meanwhile, the rapidly expandin…
Hidden in Plain Sight: Exploring Chat History Tampering in Interactive Language Models
Cheng'an Wei, Yue Zhao, Yujia Gong +3
Large Language Models (LLMs) such as ChatGPT and Llama have become prevalent in real-world applications, exhibiting impressive text generation performance. LLMs are fundamentally d…
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
Jinwen He, Yujia Gong, Kai Chen +3
Large Language Models (LLMs) have revolutionized various domains with extensive knowledge and creative capabilities. However, a critical issue with LLMs is their tendency to produc…