2 papers
stat.ML2026
An Interpretable and Scalable Framework for Evaluating Large Language Models
Xinhao Qu, Qiang Heng, Hao Zeng +1
Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM…
cs.LG2026
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
Hao Zeng, Huipeng Huang, Xinhao Qu +3
Data annotation often involves multiple sources with different cost-quality trade-offs, such as fast large language models (LLMs), slow reasoning models, and human experts. In this…