Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models
Xiao Zhang, Qianru Meng, Yongjian Chen +2
Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based…
cs.AI2025
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
Yumeng Wang, Tianyu Fan, Lingrui Xu +1
Large Language Models (LLMs) have evolved from simple chatbots into sophisticated agents capable of automating complex real-world tasks, where browsing and reasoning over live web…