307 citations · 318 across the 8 of their papers we have counts for
1 paper · 1 filter
Tingchen Fu, Mrinank Sharma, Philip Torr +3
Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBenc…