1 paper · 1 filter
Aobo Kong, Wentao Ma, Shiwan Zhao +7
Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO)…