2 papers
cs.SE2026
Automated structural testing of LLM-based agents: methods, framework, and case studies
Jens Kohl, Otto Kruse, Youssef Mostafa +9
LLM-based agents are rapidly being adopted across diverse domains. Since they interact with users without supervision, they must be tested extensively. Current testing approaches f…
cs.SE2024
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
Jens Kohl, Luisa Gloger, Rui Costa +10
As LLM-based applications reach millions of customers, ensuring their scalability and continuous quality improvement is critical for success. However, the current workflows for dev…