2 papers
cs.SE2026
CodeMind: Evaluating Large Language Models for Code Reasoning
Changshu Liu, Yang Chen, Reyhaneh Jabbarvand
Large Language Models (LLMs) have been widely used to automate programming tasks. Their capabilities have been evaluated by assessing the quality of generated code through tests or…
cs.SE2024
QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
Alex Sanchez-Stern, Abhishek Varghese, Zhanna Kaufman +3
Formal verification is a promising method for producing reliable software, but the difficulty of manually writing verification proofs severely limits its utility in practice. Recen…