8 papers
Eliciting Trustworthiness Priors of Large Language Models via Economic Games
Siyu Yan, Lusha Zhu, Jian-Qiao Zhu
One critical aspect of building human-centered, trustworthy artificial intelligence (AI) systems is maintaining calibrated trust: appropriate reliance on AI systems outperforms bot…
Simulated Annealing Enhances Theory-of-Mind Reasoning in Autoregressive Language Models
Xucong Hu, Jian-Qiao Zhu
Autoregressive language models are next-token predictors and have been criticized for only optimizing surface plausibility (i.e., local coherence) rather than maintaining correct l…
Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations
Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths
Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors." Th…
Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints
Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths
Rational decision-making under uncertainty requires coherent degrees of belief in events. However, event probabilities generated by Large Language Models (LLMs) have been shown to…
Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
Jian-Qiao Zhu, Hanbo Xie, Dilip Arumugam +2
A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural netwo…
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
Jianyu Liu, Hangyu Guo, Ranjie Duan +14
Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks…