collaborators

5 papers

cs.CV2026

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

Bingnan Liu, Chenhang Cui, Rui Huang +7

We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by…

cs.AI2026

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Utkarsh Tyagi, Xingang Guo, MohammadHossein Rezaei +5

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors requir…

cs.CL2026

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Hezhe Qiao, Hanghang Tong, Ee-Peng Lim +2

Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attri…

cs.LG2025

TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models

Rakshith S Srinivasa, Zora Che, Chen Bo Calvin Zhang +11

As students increasingly adopt large language models (LLMs) as learning aids, it is crucial to build models that are adept at handling the nuances of tutoring: they need to identif…

cs.CL2025

MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

Alexander R. Fabbri, Diego Mares, Jorge Flores +5

Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…