collaborators

7 papers

cs.AI2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Xiaomin Li, Yuexing Hao, Jianheng Hou +90

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…

cs.CL2026

User-Assistant Bias in LLMs

Xu Pan, Jingxuan Fan, Zidi Xiong +3

Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicitly mark the source of each piece…

cs.AI2026

Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning

Zidi Xiong, Shan Chen, Himabindu Lakkaraju

As Large Reasoning Models (LRMs) are increasingly deployed, auditing their chain-of-thought (CoT) traces for safety becomes critical. Recent work has reported that monitorability--…

cs.CL2026

Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments

Bingyang Ye, Shan Chen, Jingxuan Tu +4

Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientif…

cs.AI2025

How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior

Zidi Xiong, Yuping Lin, Wenya Xie +5

Memory is a critical component in large language model (LLM)-based agents, enabling them to store and retrieve past executions to improve task performance over time. In this paper,…

cs.CL2025

When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy

Jirui Qi, Shan Chen, Zidi Xiong +3

Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studi…