Position: Vibe Coding Needs Vibe Reasoning: Improving Vibe Coding with Formal Verification
arXiv:2511.00202 · doi:10.1145/3759425.3763390
Abstract
``Vibe coding'' -- the practice of developing software through iteratively conversing with a large language model (LLM) -- has exploded in popularity within the last year. However, developers report key limitations including the accumulation of technical debt, security issues, and code churn to achieve satisfactory results. We argue that these pitfalls result from LLMs' inability to reconcile accumulating human-imposed constraints during vibe coding, with developers inadvertently failing to resolve contradictions because LLMs prioritize user commands over code consistency. Given LLMs' receptiveness to verification-based feedback, we argue that formal methods can mitigate these pitfalls, making vibe coding more reliable. However, we posit that integrating formal methods must transcend existing approaches that combine formal methods and LLMs. We advocate for a side-car system throughout the vibe coding process which: (1) \emph{Autoformalizes} specifications (2) Validates against targets, (3) Delivers \emph{actionable} feedback to the LLM, and (4) Allows intuitive developer influence on specifications.
7 pages, 3 figures, In Proceedings of the 1st ACM SIGPLAN International Workshop on Language Models and Programming Languages (LMPL'25), October 12-18, 2025, Singapore, Singapore. ACM, New York, NY, USA
References in corpus (20)
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair
- LLM Security Guard for Code
- VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
- Finding Inductive Loop Invariants using Large Language Models
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
- Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
- Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability
- On the Impacts of Contexts on Repository-Level Code Generation
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
- LLM Agents Making Agent Tools
- Consistent Autoformalization for Constructing Mathematical Libraries
- FullStack Bench: Evaluating LLMs as Full Stack Coders
- Towards Autoformalization of Mathematics and Code Correctness: Experiments with Elementary Proofs
- Enhancing Automated Loop Invariant Generation for Complex Programs with Large Language Models
- FormalAlign: Automated Alignment Evaluation for Autoformalization
- Towards General Loop Invariant Generation: A Benchmark of Programs with Memory Manipulation
- Flexible and Efficient Grammar-Constrained Decoding
- Finding 709 Defects in 258 Projects: An Experience Report on Applying CodeQL to Open-Source Embedded Software (Experience Paper) -- Extended Report
- From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization