3 papers
cs.AI2026
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
Jiangyun Zhang, Kristen Surrao, Torpong Nitayanont +10
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical comput…
cs.CL2025
Self-supervised Analogical Learning using Language Models
Ben Zhou, Sarthak Jain, Yi Zhang +4
Large language models have been shown to suffer from reasoning inconsistency issues. That is, they fail more in situations unfamiliar to the training data, even though exact or ver…
cs.CL2024
BEAVER: An Enterprise Benchmark for Text-to-SQL
Peter Baile Chen, Devin Yang, Weiyue Li +6
Existing text-to-SQL benchmarks have largely been constructed from public databases with well-structured schemas and simplistic question-SQL pairs. While large language models (LLM…