From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
Debin Meng, Jiaming Yang, Zefang Zong +4
The paper introduces DataClawEval, a benchmark that tests autonomous LLM agents on end-to-end data engineering tasks across multiple production-grade SQL and Spark engines using de…
cs.AI2026
SIRIUS-SQL: Anchoring Multi-Candidate Text-to-SQL in Execution Feedback
Leo Luo, Haining Xie, Siqi Shen +8
Text-to-SQL on complex schemas is unreliable on a single pass, so recent systems generate multiple SQL candidates and let voting filter out errors. Yet voting alone is not enough,…