3 papers
cs.AI2026
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
Yongchang Peng, Qingshui Gu, Liya Zhu +31
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world…
cs.AI2026
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Liya Zhu, Xin Ma, Tao Liu +35
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on re…
cs.AI2025
LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
Liya Zhu, Peizhuang Cong, Jingzhe Ding +17
Large Language Models (LLMs) perform well on standard reasoning and question-answering benchmarks, yet such evaluations often fail to capture their ability to handle long-tail, exp…