2 citations · 7 across the 17 of their papers we have counts for
Showing 2026 · cs.CLShow all
3 papers · 2 filters
cs.CL2026
Mitigating Context Interference for Reliable and Efficient Search Agents
Boyang Xue, Bin Wu, Shuofei Qiao +8
Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts…
cs.CL2026
UXBench: Benchmarking User Experience in AI Assistants
Mengze Hong, Xia Zeng, Zeyang Lei +26
As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present UXBench, the first use…
cs.CL2026
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
Fangrui Huang, Souhad Chbeir, Arpandeep Khatua +8
Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks-…