Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian +1
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJu…
cs.AI2025
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
Bidyapati Pradhan, Surajit Dasgupta, Amit Kumar Saha +4
The advancement of large language models (LLMs) is critically dependent on the availability of high-quality datasets for Supervised Fine-Tuning (SFT), alignment tasks like Direct P…