Type-Constrained Code Generation with Language Models
arXiv:2504.09246 · doi:10.1145/3729274
Abstract
Large language models (LLMs) have achieved notable success in code generation. However, they still frequently produce uncompilable output because their next-token inference procedure does not model formal aspects of code. Although constrained decoding is a promising approach to alleviate this issue, it has only been applied to handle either domain-specific languages or syntactic features of general-purpose programming languages. However, LLMs frequently generate code with typing errors, which are beyond the domain of syntax and generally hard to adequately constrain. To address this challenge, we introduce a type-constrained decoding approach that leverages type systems to guide code generation. For this purpose, we develop novel prefix automata and a search over inhabitable types, forming a sound approach to enforce well-typedness on LLM-generated code. We formalize our approach on a foundational simply-typed language and extend it to TypeScript to demonstrate practicality. Our evaluation on the HumanEval and MBPP datasets shows that our approach reduces compilation errors by more than half and significantly increases functional correctness in code synthesis, translation, and repair tasks across LLMs of various sizes and model families, including state-of-the-art open-weight models with more than 30B parameters. The results demonstrate the generality and effectiveness of our approach in constraining LLM code generation with formal rules of type systems.
References in corpus (22)
- Code Llama: Open Foundation Models for Code
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Prompting Is Programming: A Query Language for Large Language Models
- Qwen2.5 Technical Report
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair
- A Survey of Hallucination in Large Foundation Models
- StarCoder 2 and The Stack v2: The Next Generation
- OpenAI o1 System Card
- Efficient Guided Generation for Large Language Models
- A Systematic Literature Review on Large Language Models for Automated Program Repair
- CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
- Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
- Statically Contextualizing Large Language Models with Typed Holes
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
- A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
- Syzygy: Dual Code-Test C to (safe) Rust Translation using LLMs and Dynamic Analysis
- Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
- A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
- Compilation Quotient (CQ): A Metric for the Compilation Hardness of Programming Languages
- Enhancing Code Generation for Low-Resource Languages: No Silver Bullet