belief change prediction 1dataset creation 1dialogue safety 1human-ai interaction 1language model manipulation 1
From the 1 of 13 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Xuhui Zhou, Weiwei Sun, Qianou Ma +8
As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user…
cs.AI2025
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
Zhonghao He, Tianyi Qiu, Hirokazu Shirado +1
Recent advances in reasoning techniques have substantially improved the performance of large language models (LLMs), raising expectations for their ability to provide accurate, tru…