LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems
arXiv:2401.05443 · doi:10.1145/3639477.3639743
Abstract
Although Large Language Models (LLMs) have established pre-dominance in automated code generation, they are not devoid of shortcomings. The pertinent issues primarily relate to the absence of execution guarantees for generated code, a lack of explainability, and suboptimal support for essential but niche programming languages. State-of-the-art LLMs such as GPT-4 and LLaMa2 fail to produce valid programs for Industrial Control Systems (ICS) operated by Programmable Logic Controllers (PLCs). We propose LLM4PLC, a user-guided iterative pipeline leveraging user feedback and external verification tools including grammar checkers, compilers and SMV verifiers to guide the LLM's generation. We further enhance the generation potential of LLM by employing Prompt Engineering and model fine-tuning through the creation and usage of LoRAs. We validate this system using a FischerTechnik Manufacturing TestBed (MFTB), illustrating how LLMs can evolve from generating structurally flawed code to producing verifiably correct programs for industrial applications. We run a complete test suite on GPT-3.5, GPT-4, Code Llama-7B, a fine-tuned Code Llama-7B model, Code Llama-34B, and a fine-tuned Code Llama-34B model. The proposed pipeline improved the generation success rate from 47% to 72%, and the Survey-of-Experts code quality from 2.25/10 to 7.75/10. To promote open research, we share the complete experimental setup, the LLM Fine-Tuning Weights, and the video demonstrations of the different programs on our dedicated webpage.
12 pages; 8 figures; Appearing in the 46th International Conference on Software Engineering: Software Engineering in Practice; for demo website, see https://sites.google.com/uci.edu/llm4plc/home
References in corpus (11)
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Evaluating Large Language Models Trained on Code
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
- VeriGen: A Large Language Model for Verilog Code Generation
- On Formal Reasoning on the Semantics of PLC using Coq
- The potential of LLMs for coding with low-resource and domain-specific programming languages
- Domain Adaptive Code Completion via Language Models and Decoupled Domain Databases