Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Bo Deng, Kang Zhou, Lifan Guo +6
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cov…
cs.AI2026
UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL
Jianling Gao, Chongyang Tao, Jiayuan Bai +7
Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL dialects. However, real-world…
cs.AI2025
JudgeSQL: Reasoning over SQL Candidates with Weighted Consensus Tournament
Jiayuan Bai, Xuan-guang Pan, Chongyang Tao +1
Text-to-SQL is a pivotal task that bridges natural language understanding and structured data access, yet it remains fundamentally challenging due to semantic ambiguity and complex…