2 papers
cs.AI2026
Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification
Jinghan Xu, Yikai Zhang, Aili Chen +3
Agent harnesses shape how language-model agents use instructions, tools, and runtime components, but adapting these harnesses requires costly verification. Existing propose-and-ver…
cs.CL2025
BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation
Fabian Wenz, Omar Bouattour, Devin Yang +4
Large language models (LLMs) have been successfully applied to many tasks, including text-to-SQL generation. However, much of this work has focused on publicly available datasets,…