Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems
Hezhe Qiao, Hanghang Tong, Ee-Peng Lim +2
Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attri…
cs.CL2025
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Alexander R. Fabbri, Diego Mares, Jorge Flores +5
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…