4 citations · 6 across the 14 of their papers we have counts for
1 paper · 2 filters
Kenneth Li, Samy Jelassi, Hugh Zhang +3
We present an approach called Q-probing to adapt a pre-trained language model to maximize a task-specific reward function. At a high level, Q-probing sits between heavier approache…