collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Provable Joint Decontamination for Benchmarking Multiple Large Language Models

Zhenlong Liu, Hao Zeng, Hongxin Wei

Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performa…

cs.LG2026

RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models

Sai Hao, Hao Zeng, Hongxin Wei +1

Efficiently routing queries to the optimal large language model (LLM) is crucial for optimizing the cost-performance trade-off in multi-model systems. However, most existing router…

cs.LG2026

Provable Model Provenance Set for Large Language Models

Xiaoqi Qiu, Hao Zeng, Zhiyu Hou +1

The growing prevalence of unauthorized model usage and misattribution has increased the need for reliable model provenance analysis. However, existing methods largely rely on heuri…

cs.LG2026

HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees

Hao Zeng, Huipeng Huang, Xinhao Qu +3

Data annotation often involves multiple sources with different cost-quality trade-offs, such as fast large language models (LLMs), slow reasoning models, and human experts. In this…

cs.LG2026

Distribution-informed Efficient Conformal Prediction for Full Ranking

Wenbo Liao, Huipeng Huang, Chen Jia +3

Quantifying uncertainty is critical for the safe deployment of ranking models in real-world applications. Recent work offers a rigorous solution using conformal prediction in a ful…