3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2024
Large Language Models can be Strong Self-Detoxifiers
Ching-Yun Ko, Pin-Yu Chen, Payel Das +6
Reducing the likelihood of generating harmful and toxic output is an essential task when aligning large language models (LLMs). Existing methods mainly rely on training an external…
cs.AI2021★ 3 cited
Contrastive Explanations for Comparing Preferences of Reinforcement Learning Agents
Jasmina Gajcin, Rahul Nair, Tejaswini Pedapati +3
In complex tasks where the reward function is not straightforward and consists of a set of objectives, multiple reinforcement learning (RL) policies that perform task adequately, b…