1 paper · 1 filter
Taehoon Kim, Henry Gouk, Timothy Hospedales
Test-time alignment (TTA) aims to adapt models to specific rewards during inference. However, existing methods tend to either under-optimise or over-optimise (reward hack) the targ…