4 papers
Automated structural testing of LLM-based agents: methods, framework, and case studies
Jens Kohl, Otto Kruse, Youssef Mostafa +9
LLM-based agents are rapidly being adopted across diverse domains. Since they interact with users without supervision, they must be tested extensively. Current testing approaches f…
STELLAR: A Search-Based Testing Framework for Large Language Model Applications
Lev Sorokin, Ivan Vasilev, Ken E. Friedl +1
Large Language Model (LLM)-based applications are increasingly deployed across various domains, including customer service, education, and mobility. However, these systems are pron…
Benchmarking Contextual Understanding for In-Car Conversational Systems
Philipp Habicht, Lev Sorokin, Abdullah Saydemir +2
In-Car Conversational Question Answering (ConvQA) systems significantly enhance user experience by enabling seamless voice interactions. However, assessing their accuracy and relia…
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models
Rafael Giebisch, Ken E. Friedl, Lev Sorokin +1
In-car conversational systems bring the promise to improve the in-vehicle user experience. Modern conversational systems are based on Large Language Models (LLMs), which makes them…