3 papers
cs.SE2026
SONAR: Task-Aware Code Summary Evaluation for LLM Consumers Without References
Simantika Bhattacharjee Dristi, Matthew B. Dwyer
Source code summaries have traditionally been evaluated from a human developer's perspective, with quality determined by how closely they resemble developer-written references and…
cs.SE2026
A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings
Simantika Bhattacharjee Dristi, Matthew B. Dwyer
With the rapid adoption of large language models (LLMs) in automated code refactoring, assessing and ensuring functional equivalence between LLM-generated refactoring and the origi…
cs.SE2025
Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
Simantika Bhattacharjee Dristi, Matthew B. Dwyer
With the increasing popularity of large language models (LLMs) and LLM-based agents, reliable and effective code evaluation metrics (CEMs) have become crucial for progress across s…