1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Ayoub Hammal, Pierre Zweigenbaum, Caio Corro
Recent works proposed test-time alignment methods that rely on a small aligned model as a proxy that guides the generation of a larger base (unaligned) model. The implicit reward a…