4 papers
CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues
Guoxiang Guo, Kla Tantithamthavorn, Neelofar Neelofar +2
Large Language Models (LLMs) are increasingly used in software engineering to generate and refine code. In practice, developers often continue from an initial code generation reque…
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
Aaron Guoxiang Guo, Aldeida Aleti, Neelofar Neelofar +3
With the widespread application of LLM-based dialogue systems in daily life, quality assurance has become more important than ever. Recent research has successfully introduced meth…
UntrustVul: An Automated Approach for Identifying Untrustworthy Alerts in Vulnerability Detection Models
Lam Nguyen Tung, Xiaoning Du, Neelofar Neelofar +1
Machine learning (ML) has shown promise in vulnerability detection, but ML detectors may rely on irrelevant code features, causing them to highlight non-vulnerable lines as suspici…
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
Lam Nguyen Tung, Steven Cho, Xiaoning Du +4
Machine learning (ML) for text classification has been widely used in various domains. These applications can significantly impact ethics, economics, and human behavior, raising se…