11 papers
On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study
Daniel Rodriguez-Cardenas, David Nader Palacio, Anna Schmedding +8
Modern software systems require earlier and more scalable vulnerability severity assessment to reduce exposure to high-impact security flaws. Security analysts typically assign CVS…
ECLAIR: A Causally-Grounded AI Framework for Scientific Discovery in Empirical Software Engineering
Alejandro Velasco, Daniel Rodriguez-Cardenas, Dipin Khati +2
The scientific method has long guided empirical research in Software Engineering (SE), but the complexity of modern software systems often hinders its systematic application. This…
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
Nilesh Nayan, Aishwarya Sampath Kumar, Rishiraj Girmal +5
Safety benchmarks assume that test-condition behavior predicts deployment behavior, an assumption that fails if models detect evaluation cues and adapt. This opens a gap between be…
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
Dipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio +3
As Large Language Models for Code (LM4Code) become integral to software engineering, establishing trust in their output becomes critical. However, standard accuracy metrics obscure…
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
Alejandro Velasco, Daniel Rodriguez-Cardenas, Dipin Khati +3
Recent advances in large language models (LLMs) have accelerated their adoption in software engineering contexts. However, concerns persist about the structural quality of the code…
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
Dipin Khati, Yijin Liu, David N. Palacio +2
Applications of Large Language Models (LLMs) are rapidly growing in industry and academia for various software engineering (SE) tasks. As these models become more integral to criti…