4 papers
Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
Hao Yang, Haoxuan Li, Luyu Chen +3
Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved variables can lead to biased causal estimates. While recent wo…
RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems
Luyu Chen, Quanyu Dai, Zeyu Zhang +6
Conversational recommender systems (CRS) enhance user experience through multi-turn interactions, yet evaluating CRS remains challenging. User simulators can provide comprehensive…
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
Luyu Chen, Zeyu Zhang, Haoran Tan +4
LLMs have emerged as powerful evaluators in the LLM-as-a-Judge paradigm, offering significant efficiency and flexibility compared to human judgments. However, previous methods prim…
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
Zeyu Zhang, Quanyu Dai, Luyu Chen +7
LLM-based agents have been widely applied as personal assistants, capable of memorizing information from user messages and responding to personal queries. However, there still lack…