Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
When Bad Data Leads to Good Models
Kenneth Li, Yida Chen, Fernanda Viégas +1
In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- an…
cs.LG2024
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
Kenneth Li, Samy Jelassi, Hugh Zhang +3
We present an approach called Q-probing to adapt a pre-trained language model to maximize a task-specific reward function. At a high level, Q-probing sits between heavier approache…