5 papers
Improving Methodologies for LLM Evaluations Across Global Languages
Akriti Vij, Benjamin Chua, Darshini Ramiah +43
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current…
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
Boming Xia, Qinghua Lu, Liming Zhu +3
Large Language Models (LLMs) have enabled the emergence of LLM agents, systems capable of pursuing under-specified goals and adapting after deployment. Evaluating such agents is ch…
AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
Qinghua Lu, Dehai Zhao, Yue Liu +6
The emergence of foundation models (FMs) has enabled the development of highly capable and autonomous agents, unlocking new application opportunities across a wide range of domains…
DOCUEVAL: An LLM-based AI Engineering Tool for Building Customisable Document Evaluation Workflows
Hao Zhang, Qinghua Lu, Liming Zhu
Foundation models, such as large language models (LLMs), have the potential to streamline evaluation workflows and improve their performance. However, practical adoption faces chal…