Fully Autonomous Programming using Iterative Multi-Agent Debugging with Large Language Models
arXiv:2503.07693 · doi:10.1145/3719351
Abstract
Program synthesis with Large Language Models (LLMs) suffers from a "near-miss syndrome": the generated code closely resembles a correct solution but fails unit tests due to minor errors. We address this with a multi-agent framework called Synthesize, Execute, Instruct, Debug, and Repair (SEIDR). Effectively applying SEIDR to instruction-tuned LLMs requires determining (a) optimal prompts for LLMs, (b) what ranking algorithm selects the best programs in debugging rounds, and (c) balancing the repair of unsuccessful programs with the generation of new ones. We empirically explore these trade-offs by comparing replace-focused, repair-focused, and hybrid debug strategies. We also evaluate lexicase and tournament selection to rank candidates in each generation. On Program Synthesis Benchmark 2 (PSB2), our framework outperforms both conventional use of OpenAI Codex without a repair phase and traditional genetic programming approaches. SEIDR outperforms the use of an LLM alone, solving 18 problems in C++ and 20 in Python on PSB2 at least once across experiments. To assess generalizability, we employ GPT-3.5 and Llama 3 on the PSB2 and HumanEval-X benchmarks. Although SEIDR with these models does not surpass current state-of-the-art methods on the Python benchmarks, the results on HumanEval-C++ are promising. SEIDR with Llama 3-8B achieves an average pass@100 of 84.2%. Across all SEIDR runs, 163 of 164 problems are solved at least once with GPT-3.5 in HumanEval-C++, and 162 of 164 with the smaller Llama 3-8B. We conclude that SEIDR effectively overcomes the near-miss syndrome in program synthesis with LLMs.
Accepted for publication in ACM Trans. Evol. Learn. Optim., February 2025. arXiv admin note: text overlap with arXiv:2304.10423
References in corpus (26)
- LLaMA: Open and Efficient Foundation Language Models
- A Survey of Deep Learning Techniques for Autonomous Driving
- Evaluating Large Language Models Trained on Code
- Code Llama: Open Foundation Models for Code
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Random Walks: A Review of Algorithms and Applications
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Mining Idioms from Source Code
- Teaching Large Language Models to Self-Debug
- Code Search: A Survey of Techniques for Finding Code
- Fully Autonomous Programming with Large Language Models
- Conversational Automated Program Repair
- Self-collaboration Code Generation via ChatGPT
- Towards Better Chain-of-Thought Prompting Strategies: A Survey
- SelfEvolve: A Code Evolution Framework via Large Language Models
- A Systematic Literature Review on Large Language Models for Automated Program Repair
- Algorithm Evolution Using Large Language Model
- Repair Is Nearly Generation: Multilingual Program Repair with LLMs
- Synthesize, Execute and Debug: Learning to Repair for Neural Program Synthesis
- Automated Repair of Programs from Large Language Models
- Exploring the Robustness of Large Language Models for Solving Programming Problems
- Evolution through Large Models
- CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code Models
- PSB2: The Second Program Synthesis Benchmark Suite
- Autoencoders as Tools for Program Synthesis