Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa +8
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…
cs.AI2026
Collaborative Belief Reasoning with LLMs for Efficient Multi-Agent Collaboration
Zhimin Wang, Duo Wu, Shaokang He +6
Effective real-world multi-agent collaboration requires not only accurate planning but also the ability to reason about collaborators' intents--a crucial capability for avoiding mi…