238 citations · 250 across the 10 of their papers we have counts for
3 papers · 1 filter
BOND: Aligning LLMs with Best-of-N Distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17
Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-t…
Zeroth-order non-convex learning via hierarchical dual averaging
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization - i.e., learning processes where, at each stage, the optimizer is facing an unkn…
Online non-convex optimization with imperfect feedback
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes - or otherwise constructs - an inexact model for the lo…