1 paper · 1 filter
Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull +4
Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to rewa…