Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian +1
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJu…
cs.AI2025
FABRIC: Framework for Agent-Based Realistic Intelligence Creation
Abhigya Verma, Seganrasan Subramanian, Nandhakumar Kandasamy +1
Large language models (LLMs) are increasingly deployed as agents, expected to decompose goals, invoke tools, and verify results in dynamic environments. Realizing these capabilitie…
cs.AI2025
GRAFT: GRaPH and Table Reasoning for Textual Alignment -- A Benchmark for Structured Instruction Following and Visual Reasoning
Abhigya Verma, Sriram Puttagunta, Seganrasan Subramanian +1
GRAFT is a structured multimodal benchmark designed to probe how well LLMs handle instruction following, visual reasoning, and tasks requiring tight visual textual alignment. The d…