2 papers
cs.LG2026
Conditional Evaluation of Language Models with Cheap Auxiliary Signals
Zhi Zhang, Lingfeng Lyu, Yue Kang +1
Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-j…
stat.ML2026
Inferential Evaluation of Surrogate-Derived Models under Covariate Shift
Longtian Shi, Molei Liu, Doudou Zhou
In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved. Evaluating its tar…