3 papers
cs.LG2025
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Tzu-Heng Huang, Harit Vishwakarma, Frederic Sala
Large language models (LLMs) are widely used to evaluate the quality of LLM generations and responses, but this leads to significant challenges: high API costs, uncertain reliabili…
cs.LG2025
ScriptoriumWS: A Code Generation Assistant for Weak Supervision
Tzu-Heng Huang, Catherine Cao, Spencer Schoenberg +3
Weak supervision is a popular framework for overcoming the labeled data bottleneck: the need to obtain labels for training data. In weak supervision, multiple noisy-but-cheap sourc…
cs.LG2024
OTTER: Effortless Label Distribution Adaptation of Zero-shot Models
Changho Shin, Jitian Zhao, Sonia Cromp +2
Popular zero-shot models suffer due to artifacts inherited from pretraining. One particularly detrimental issue, caused by unbalanced web-scale pretraining data, is mismatched labe…