Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
Yongchang Peng, Qingshui Gu, Liya Zhu +31
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world…
cs.AI2026
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Liya Zhu, Xin Ma, Tao Liu +35
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on re…