1 paper
Yao Liu, Guangjia Chai, Yuming Huang +3
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate…