Mokav: Execution-driven Differential Testing with LLMs
arXiv:2406.10375 · doi:10.1016/j.jss.2025.112571
Abstract
It is essential to detect functional differences between programs in various software engineering tasks, such as automated program repair, mutation testing, and code refactoring. The problem of detecting functional differences between two programs can be reduced to searching for a difference exposing test (DET): a test input that results in different outputs on the subject programs. In this paper, we propose Mokav, a novel execution-driven tool that leverages LLMs to generate DETs. Mokav takes two versions of a program (P and Q) and an example test input. When successful, Mokav generates a valid DET, a test input that leads to provably different outputs on P and Q. Mokav iteratively prompts an LLM with a specialized prompt to generate new test inputs. At each iteration, Mokav provides execution-based feedback from previously generated tests until the LLM produces a DET. We evaluate Mokav on 1535 pairs of Python programs collected from the Codeforces competition platform and 32 pairs of programs from the QuixBugs dataset. Our experiments show that Mokav outperforms the state-of-the-art, Pynguin and Differential Prompting, by a large margin. Mokav can generate DETs for 81.7% (1,255/1535) of the program pairs in our benchmark (versus 4.9% for Pynguin and 37.3% for Differential Prompting). We demonstrate that the iterative and execution-driven feedback components of the system contribute to its high effectiveness.
References in corpus (24)
- Evaluating Large Language Models Trained on Code
- DLFuzz: Differential Fuzzing Testing of Deep Learning Systems
- Fuzz4All: Universal Fuzzing with Large Language Models
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Measuring Coding Challenge Competence With APPS
- TOGA: A Neural Method for Test Oracle Generation
- On Learning Meaningful Assert Statements for Unit Test Cases
- Generating Accurate Assert Statements for Unit Test Cases using Pretrained Transformers
- Pynguin: Automated Unit Test Generation for Python
- CodeT: Code Generation with Generated Tests
- An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation
- Traces of Memorisation in Large Language Models for Code
- ChatUniTest: A Framework for LLM-Based Test Generation
- Interactive Code Generation via Test-Driven User-Intent Formalization
- TOGLL: Correct and Strong Test Oracle Generation with LLMs
- Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation
- Can Large Language Models Write Good Property-Based Tests?
- PyTester: Deep Reinforcement Learning for Text-to-Testcase Generation
- Large Language Models as Test Case Generators: Performance Evaluation and Enhancement
- Augmenting Diffs With Runtime Information
- Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
- exLong: Generating Exceptional Behavior Tests with Large Language Models
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
- Multi-Programming Language Sandbox for LLMs