activity
20242026
collaborators

8 papers

cs.CL2026

Eliciting Trustworthiness Priors of Large Language Models via Economic Games

Siyu Yan, Lusha Zhu, Jian-Qiao Zhu

One critical aspect of building human-centered, trustworthy artificial intelligence (AI) systems is maintaining calibrated trust: appropriate reliance on AI systems outperforms bot…

cs.CL2026

Simulated Annealing Enhances Theory-of-Mind Reasoning in Autoregressive Language Models

Xucong Hu, Jian-Qiao Zhu

Autoregressive language models are next-token predictors and have been criticized for only optimizing surface plausibility (i.e., local coherence) rather than maintaining correct l…

cs.CL2025

Steering Risk Preferences in Large Language Models by Aligning Behavioral and Neural Representations

Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths

Changing the behavior of large language models (LLMs) can be as straightforward as editing the Transformer's residual streams using appropriately constructed "steering vectors." Th…

cs.CL2025

Recovering Event Probabilities from Large Language Model Embeddings via Axiomatic Constraints

Jian-Qiao Zhu, Haijiang Yan, Thomas L. Griffiths

Rational decision-making under uncertainty requires coherent degrees of belief in events. However, event probabilities generated by Large Language Models (LLMs) have been shown to…

cs.AI2025

Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions

Jian-Qiao Zhu, Hanbo Xie, Dilip Arumugam +2

A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural netwo…

cs.CL2025

DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models

Jianyu Liu, Hangyu Guo, Ranjie Duan +14

Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks…