1 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 1 cited
Is Elo Rating Reliable? A Study Under Model Misspecification
Shange Tang, Yuanhao Wang, Chi Jin
Elo rating, widely used for skill assessment across diverse domains ranging from competitive games to large language models, is often understood as an incremental update algorithm…
cs.LG2025★ 1 cited
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Kaixuan Huang, Jiacheng Guo, Zihao Li +15
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieve…
stat.ML2023★ 1 cited
Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift
Jiawei Ge, Shange Tang, Jianqing Fan +2
A key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization -- generalizing to target data whose distribution differs from that of sou…