2 papers
stat.ML2026
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
Zihan Dong, Xiaotian Hou, Ruijia Wu +1
The increasing reliance on human preference feedback to judge AI-generated pseudo labels has created a pressing need for principled, budget-conscious data acquisition strategies. W…
cs.LG2026
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
Zihan Dong, Zhixian Zhang, Yang Zhou +3
Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable ranking…