1 citations · 1 across the 3 of their papers we have counts for
4 papers
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
Daniel Rodriguez-Cardenas, Xiaochang Li, Marcos Macedo +5
Large language models for code are advancing fast, yet our ability to evaluate them lags behind. Current benchmarks focus on narrow tasks and single metrics, which hide critical ga…
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
Dipin Khati, Daniel Rodriguez-Cardenas, Paul Pantzer +1
Large Language Models (LLMs) for code generation boost productivity but frequently introduce Knowledge Conflicting Hallucinations (KCHs), subtle, semantic errors, such as non-exist…
Tricky: Towards a Benchmark for Evaluating Human and LLM Error Interactions
Cole Granger, Dipin Khati, Daniel Rodriguez-Cardenas +1
Large language models (LLMs) are increasingly integrated into software development workflows, yet they often introduce subtle logic or data-misuse errors that differ from human bug…
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
Dipin Khati, Yijin Liu, David N. Palacio +2
Applications of Large Language Models (LLMs) are rapidly growing in industry and academia for various software engineering (SE) tasks. As these models become more integral to criti…