2 papers
cs.AI2026
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont +10
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical comput…
cs.LG2026
PACE: Two-Timescale Self-Evolution for Small Language Model Agents
Chen Ling, Pei Chen, Albert Guan +4
Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline.…