2 citations · 2 across the 2 of their papers we have counts for
4 papers
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
Nazanin Mohammadi Sepahvand, Eleni Triantafillou, Hugo Larochelle +3
Large language models (LLMs) trained on webscale data can produce toxic outputs, raising concerns for safe deployment. Prior defenses, based on applications of DPO, NPO, and simila…
ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
Gaurav Sahu, Hugo Larochelle, Laurent Charlin +1
Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular…
The Search for Squawk: Agile Modeling in Bioacoustics
Vincent Dumoulin, Otilia Stretcu, Jenny Hamer +16
Passive acoustic monitoring (PAM) has shown great promise in helping ecologists understand the health of animal populations and ecosystems. However, extracting insights from millio…
Capturing Individual Human Preferences with Reward Features
André Barreto, Vincent Dumoulin, Yiran Mao +6
Reinforcement learning from human feedback usually models preferences using a reward function that does not distinguish between people. We argue that this is unlikely to be a good…