1 paper
Yuming, Huang, Yao Liu +3
Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional suppo…