5 papers · 1 filter
Provable Joint Decontamination for Benchmarking Multiple Large Language Models
Zhenlong Liu, Hao Zeng, Hongxin Wei
Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performa…
RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
Sai Hao, Hao Zeng, Hongxin Wei +1
Efficiently routing queries to the optimal large language model (LLM) is crucial for optimizing the cost-performance trade-off in multi-model systems. However, most existing router…
Provable Model Provenance Set for Large Language Models
Xiaoqi Qiu, Hao Zeng, Zhiyu Hou +1
The growing prevalence of unauthorized model usage and misattribution has increased the need for reliable model provenance analysis. However, existing methods largely rely on heuri…
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
Hao Zeng, Huipeng Huang, Xinhao Qu +3
Data annotation often involves multiple sources with different cost-quality trade-offs, such as fast large language models (LLMs), slow reasoning models, and human experts. In this…
Distribution-informed Efficient Conformal Prediction for Full Ranking
Wenbo Liao, Huipeng Huang, Chen Jia +3
Quantifying uncertainty is critical for the safe deployment of ranking models in real-world applications. Recent work offers a rigorous solution using conformal prediction in a ful…