3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.SE2024★ 3 cited
Project MPG: towards a generalized performance benchmark for LLM capabilities
Lucas Spangher, Tianle Li, William F. Arnold +6
There exists an extremely wide array of LLM benchmarking tasks, whereas oftentimes a single number is the most actionable for decision-making, especially by non-experts. No such ag…
cs.CL2024
Improving Multi-Agent Debate with Sparse Communication Topology
Yunxuan Li, Yibing Du, Jiageng Zhang +4
Multi-agent debate has proven effective in improving large language models quality for reasoning and factuality tasks. While various role-playing strategies in multi-agent debates…