Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation
Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull +4
Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to rewa…
cs.CL2025
DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
Ziyi Zhu, Olivier Tieleman, Caitlin A. Stamatis +5
Realistic user simulation is crucial for training and evaluating multi-turn dialogue systems, yet creating simulators that accurately replicate human behavior remains a significant…