2 papers
cs.SE2026
ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation
Yiru Dong, Richong Zhang, Fanshuang Kong +1
Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementa…
cs.MA2026
CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents
Jiarun Fu, Lizhong Ding, Sida Chen +6
Agent harnesses have become the operational infrastructure of modern large language model agents, coordinating context, tools, verification, and execution control to translate late…