37 papers
TARS: A Theory-of-Mind Agent for Personalized In-IDE Code Comprehension
Leopoldo Todisco, Antonio Della Porta, Stefano Lambiase +1
Code comprehension is one of the most time-consuming tasks in software engineering, yet most LLM-based assistants produce explanations that ignore who is asking and force developer…
The Language of Security: How Prompt Syntax Shapes Secure Code Generation in Open LLMs
Matteo Cicalese, Antonio Della Porta, Stefano Lambiase +4
Large Language Models (LLMs) are increasingly used for source code generation despite their outputs often exhibiting security vulnerabilities. Prior work shows that prompt engineer…
How Do Developers Maintain and Evolve Their Agents' Instructions? An Empirical Study
Gianmario Voria, Alfonso Cannavale, Andrea De Lucia +3
Context. Autonomous coding agents are increasingly used in software development, shifting parts of the engineering process to AI assistance. While this automation brings clear bene…
An Empirical Study of Gemini 3 for Detecting Natural Language Test Smells in Manual Test Cases
Keila Lucas, Rohit Gheyi, Márcio Ribeiro +4
Manual testing, in which testers follow natural language instructions to validate system behavior, remains essential for uncovering issues that are difficult to capture with automa…
LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda
Victoria Gomes, Delaney Selb, Fabio Palomba +2
Context: Empirical Software Engineering (ESE) faces increasing challenges due to data scale, methodological complexity, and reproducibility concerns. Large Language Models (LLMs) h…
From Pixels to Explanations: Interpretable Diabetic Retinopathy Grading with CNN-Transformer Ensembles, Visual Explainability and Vision-Language Models
Pir Bakhsh Khokhar, Carmine Gravino, Fabio Palomba +2
The quality of diabetic retinopathy (DR) screening relies on the ability to correctly grade severity; however, many deep-learning (DL) classifiers cannot be easily interpreted in t…