most citedDetecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.SEShow all

5 papers · 1 filter

cs.SE2026

Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering

Daniel Rodriguez-Cardenas, Xiaochang Li, Marcos Macedo +5

Large language models for code are advancing fast, yet our ability to evaluate them lags behind. Current benchmarks focus on narrow tasks and single metrics, which hide critical ga…

cs.SE20261 cited

Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis

Dipin Khati, Daniel Rodriguez-Cardenas, Paul Pantzer +1

Large Language Models (LLMs) for code generation boost productivity but frequently introduce Knowledge Conflicting Hallucinations (KCHs), subtle, semantic errors, such as non-exist…

cs.SE2026

Tricky: Towards a Benchmark for Evaluating Human and LLM Error Interactions

Cole Granger, Dipin Khati, Daniel Rodriguez-Cardenas +1

Large language models (LLMs) are increasingly integrated into software development workflows, yet they often introduce subtle logic or data-misuse errors that differ from human bug…

cs.SE2025

Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation

Dipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio +3

As Large Language Models for Code (LM4Code) become integral to software engineering, establishing trust in their output becomes critical. However, standard accuracy metrics obscure…

cs.SE2025

Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives

Dipin Khati, Yijin Liu, David N. Palacio +2

Applications of Large Language Models (LLMs) are rapidly growing in industry and academia for various software engineering (SE) tasks. As these models become more integral to criti…