5 papers
CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues
Guoxiang Guo, Kla Tantithamthavorn, Neelofar Neelofar +2
Large Language Models (LLMs) are increasingly used in software engineering to generate and refine code. In practice, developers often continue from an initial code generation reque…
UntrustVul: An Automated Approach for Identifying Untrustworthy Alerts in Vulnerability Detection Models
Lam Nguyen Tung, Xiaoning Du, Neelofar Neelofar +1
Machine learning (ML) has shown promise in vulnerability detection, but ML detectors may rely on irrelevant code features, causing them to highlight non-vulnerable lines as suspici…
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
Aaron Guoxiang Guo, Aldeida Aleti, Neelofar Neelofar +3
With the widespread application of LLM-based dialogue systems in daily life, quality assurance has become more important than ever. Recent research has successfully introduced meth…
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
Lam Nguyen Tung, Steven Cho, Xiaoning Du +4
Machine learning (ML) for text classification has been widely used in various domains. These applications can significantly impact ethics, economics, and human behavior, raising se…
Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of System-level Testing of Autonomous Vehicles
Neelofar Neelofar, Aldeida Aleti
AI-powered systems have gained widespread popularity in various domains, including Autonomous Vehicles (AVs). However, ensuring their reliability and safety is challenging due to t…