5 papers · 1 filter
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
Tzu-Heng Huang, Harit Vishwakarma, Frederic Sala
Large language models (LLMs) are widely used to evaluate the quality of LLM generations and responses, but this leads to significant challenges: high API costs, uncertain reliabili…
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
Albert Ge, Tzu-Heng Huang, John Cooper +7
Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermi…
Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis
Sonia Cromp, Satya Sai Srinath Namburi GNVV, Mohammed Alkhudhayri +4
While advances in large language models (LLMs) have greatly improved the quality of synthetic text data in recent years, synthesizing tabular data has received relatively less atte…
ScriptoriumWS: A Code Generation Assistant for Weak Supervision
Tzu-Heng Huang, Catherine Cao, Spencer Schoenberg +3
Weak supervision is a popular framework for overcoming the labeled data bottleneck: the need to obtain labels for training data. In weak supervision, multiple noisy-but-cheap sourc…
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks
Tianyi Zhang, Linrong Cai, Jeffrey Li +4
Weak supervision (WS) is a popular approach for label-efficient learning, leveraging diverse sources of noisy but inexpensive weak labels to automatically annotate training data. D…