4 papers
Tail-Aware Top- On-Policy Distillation
Huipeng Huang, Hongxin Wei
On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token distributio…
Model-agnostic Selective Labeling with Provable Statistical Guarantees
Huipeng Huang, Wenbo Liao, Huajun Xi +3
Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting la…
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
Hao Zeng, Huipeng Huang, Xinhao Qu +3
Data annotation often involves multiple sources with different cost-quality trade-offs, such as fast large language models (LLMs), slow reasoning models, and human experts. In this…
Distribution-informed Efficient Conformal Prediction for Full Ranking
Wenbo Liao, Huipeng Huang, Chen Jia +3
Quantifying uncertainty is critical for the safe deployment of ranking models in real-world applications. Recent work offers a rigorous solution using conformal prediction in a ful…