2 papers
cs.CL2026
An Empirical Study of Automating Agent Evaluation
Kang Zhou, Sangmin Woo, Haibo Ding +14
Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises:…
cs.CL2024
DoPAMine: Domain-specific Pre-training Adaptation from seed-guided data Mining
Vinayak Arannil, Neha Narwal, Sourav Sanjukta Bhabesh +5
Large Language Models (LLMs) have shown remarkable ability to generalize effectively across numerous industry domains while executing a range of tasks. Many of these competencies a…