4 papers · 1 filter
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths
Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors." Th…
Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints
Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths
Rational decision-making under uncertainty requires coherent degrees of belief in events. However, event probabilities generated by Large Language Models (LLMs) have been shown to…
Identifying and Mitigating the Influence of the Prior Distribution in Large Language Models
Liyi Zhang, Veniamin Veselovsky, R. Thomas McCoy +1
Large language models (LLMs) sometimes fail to respond appropriately to deterministic tasks -- such as counting or forming acronyms -- because the implicit prior distribution they…
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
Veniamin Veselovsky, Berke Argin, Benedikt Stroebl +5
Just as humans display language patterns influenced by their native tongue when speaking new languages, LLMs often default to English-centric responses even when generating in othe…