Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
OPD+: Rethinking the Advantage Design for On-Policy Distillation
Hanyang Zhao, Haoxian Chen, Han Lin +3
On-policy distillation (OPD) is a widely used technique to transfer capabilities from capable teacher language models to the base student models, and can be formulated in a reinfor…
cs.LG2025
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
Genta Indra Winata, David Anugraha, Emmy Liu +17
High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challe…