collaborators

8 papers

cs.CL2026

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

Tiancheng Hu, Benjamin Minixhofer, Nigel Collier

The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliab…

cs.CL2026

Multi-agent AI systems outperform human teams in creativity

Tiancheng Hu, Yixuan Jiang, Haotian Li +5

Although artificial intelligence (AI) now matches or exceeds human performance across numerous cognitive tasks, creativity remains a highly contested frontier. As AI systems based…

cs.CL2026

SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors

Tiancheng Hu, Joachim Baumann, Lorenzo Lupo +3

Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human b…

cs.CL2026

Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores

Esma Balkır, Alice Pernthaller, Marco Basaldella +2

Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tas…

cs.CL2026

Value of Information: A Framework for Human-Agent Communication

Yijiang River Dong, Tiancheng Hu, Zheng Hui +4

Large Language Model (LLM) agents deployed for real-world tasks face a fundamental dilemma: user requests are underspecified, yet agents must decide whether to act on incomplete in…

cs.CL2026

Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding

Yijiang River Dong, Tiancheng Hu, Zheng Hui +1

Large language models excel at complex instructions yet struggle to deviate from their helpful assistant persona, as post-training instills strong priors that resist conflicting in…