2 papers
cs.SE2026
When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling
Niful Islam, Ragib Shahriar Ayon, Deepak George Thomas +2
Large Language Models (LLMs) have revolutionized intelligent application development. While standalone LLMs cannot perform any actions, LLM agents address the limitation by integra…
cs.SE2024
muPRL: A Mutation Testing Pipeline for Deep Reinforcement Learning based on Real Faults
Deepak-George Thomas, Matteo Biagiola, Nargiz Humbatova +4
Reinforcement Learning (RL) is increasingly adopted to train agents that can deal with complex sequential tasks, such as driving an autonomous vehicle or controlling a humanoid rob…