3 papers
cs.AI2026
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian +1
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJu…
cs.CV2026
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
Abhigya Verma, Khyati Mahajan, Amit Kumar Saha +4
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documen…
cs.AI2025
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
Bidyapati Pradhan, Surajit Dasgupta, Amit Kumar Saha +4
The advancement of large language models (LLMs) is critically dependent on the availability of high-quality datasets for Supervised Fine-Tuning (SFT), alignment tasks like Direct P…