1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Javier Rando, Florian Tramèr
Reinforcement Learning from Human Feedback (RLHF) is used to align large language models to produce helpful and harmless responses. Yet, prior work showed these models can be jailb…