13 papers
Requirements-Augmented Generation for Trustworthy Acceptance Testing of LLM-Based Software
Fanyu Wang, Chetan Arora, Zhenping Xie +4
LLM-based software (LBS) integrates large language models as core components to deliver flexible, personalised responses. Unlike traditional software with deterministic outputs, LB…
Towards a Risk Assessment of Malicious Skill Files in Coding Agents
Rui Yang, Michael Fu, Kla Tantithamthavorn +2
Autonomous coding agents are increasingly embedded in enterprise software workflows with delegated authority over connected systems. Central to this architecture is the agent skill…
SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
Rui Yang, Michael Fu, Kla Tantithamthavorn +2
Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing domain-specific skill files to tailor age…
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
Sebastian Baltes, Florian Angermeir, Chetan Arora +19
Large Language Models (LLMs) are widely used in software engineering (SE) research and practice, yet their non-determinism, opaque training data, and rapidly evolving models threat…
Reporting LLM Prompting in Automated Software Engineering: A Guideline Based on Current Practices and Expectations
Alexander Korn, Lea Zaruchas, Chetan Arora +4
Large Language Models, particularly decoder-only generative models such as GPT, are increasingly used to automate Software Engineering tasks. These models are primarily guided thro…
DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
Rui Yang, Michael Fu, Chakkrit Tantithamthavorn +3
Intelligent software systems powered by Large Language Models (LLMs) are increasingly deployed in critical sectors, raising concerns about their safety during runtime. Through an i…