2 papers
cs.CL2026
Toward Autonomous Long-Horizon Engineering for ML Research
Guoxin Chen, Jie Chen, Lei Chen +7
Agentic systems increasingly automate pieces of AI research. Yet turning underspecified research objectives into runnable, experimentally validated ML systems remains a central bot…
cs.DC2026
ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness
Wenxing Zhu, Simeng Qi, Junkui Chen +7
We present ACE-Bench (Azure SDK Coding Evaluation Benchmark), an execution-free benchmark that provides fast, reproducible pass or fail signals for whether large language model (LL…