2 papers
cs.AI2026
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
Yongchang Peng, Qingshui Gu, Liya Zhu +31
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world…
cs.AI2026
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
Xue Liu, Xin Ma, Yuxin Ma +36
As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks c…