Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
Jessica Hoffmann, Christiane Ahlheim, Zac Yu +8
The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on…
cs.CL2024
Gradient-Based Language Model Red Teaming
Nevan Wichers, Carson Denison, Ahmad Beirami
Red teaming is a common strategy for identifying weaknesses in generative language models (LMs), where adversarial prompts are produced that trigger an LM to generate unsafe respon…