5 papers
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
Bingnan Liu, Chenhang Cui, Rui Huang +7
We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by…
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei +5
Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors requir…
VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems
Hezhe Qiao, Hanghang Tong, Ee-Peng Lim +2
Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attri…
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
Rakshith S Srinivasa, Zora Che, Chen Bo Calvin Zhang +11
As students increasingly adopt large language models (LLMs) as learning aids, it is crucial to build models that are adept at handling the nuances of tutoring: they need to identif…
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Alexander R. Fabbri, Diego Mares, Jorge Flores +5
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…