3 papers
cs.AI2026
Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes
Liang Chen
LLM-as-user-simulation has become core infrastructure for conversational AI: agent benchmarks (tau-bench), training pipelines, and a growing body of fidelity studies all rely on LL…
cs.CL2026
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
Liang Chen, Qi Liu, Wenhuan Lin +1
Multi-dimensional rubric-based dialogue evaluation is widely used to assess conversational AI, yet its criterion validity -- whether quality scores are associated with the downstre…
cs.LG2026
Neural Paging: Learning Context Management Policies for Turing-Complete Agents
Liang Chen, Qi Liu
The proof that Large Language Models (LLMs) augmented with external read-write memory constitute a computationally universal system has established the theoretical foundation for g…