2 papers
cs.SE2026
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Yuangang Li, Justin Tian Jin Chen, Ethan Yu +2
Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning eva…
cs.SE2025
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?
QiHong Chen, Jiachen Yu, Jiawei Li +3
Recent advancements in Large Language Models (LLMs) have led to their widespread application in automated code generation. However, these models can still generate defective code t…