13 papers
Provable Model Provenance Set for Large Language Models
Xiaoqi Qiu, Hao Zeng, Zhiyu Hou +1
The growing prevalence of unauthorized model usage and misattribution has increased the need for reliable model provenance analysis. However, existing methods largely rely on heuri…
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
Hao Zeng, Huipeng Huang, Xinhao Qu +3
Data annotation often involves multiple sources with different cost-quality trade-offs, such as fast large language models (LLMs), slow reasoning models, and human experts. In this…
Distribution-informed Efficient Conformal Prediction for Full Ranking
Wenbo Liao, Huipeng Huang, Chen Jia +3
Quantifying uncertainty is critical for the safe deployment of ranking models in real-world applications. Recent work offers a rigorous solution using conformal prediction in a ful…
Conditional Performance Guarantee for Large Reasoning Models
Jianguo Huang, Hao Zeng, Bingyi Jing +2
Large reasoning models have shown strong performance through extended chain-of-thought reasoning, yet their computational cost remains significant. Probably approximately correct (…
Multi-Condition Conformal Selection
Qingyang Hao, Wenbo Liao, Bingyi Jing +1
Selecting high-quality candidates from large-scale datasets is critically important in resource-constrained applications such as drug discovery, precision medicine, and the alignme…
Model-agnostic Selective Labeling with Provable Statistical Guarantees
Huipeng Huang, Wenbo Liao, Huajun Xi +3
Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting la…