2 papers
cs.CL2026
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
Yao Liu, Guangjia Chai, Yuming Huang +3
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate…
cs.CL2026
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm
Yuming, Huang, Yao Liu +3
Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional suppo…