2 papers
stat.ML2026
A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth
Mingyuan Xu, Xinzi Tan, Jiawei Wu +1
Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is…
stat.ME2026
Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes
Mingyuan Xu, Zongqi Xia, Tianxi Cai +2
We often collect data from multiple sites (e.g., hospitals) that share common structure but also exhibit heterogeneity. This paper aims to learn robust sequential decision-making p…