3 papers
cs.AI2025
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
Dilani Widanapathiranage, Scott Barnett, Stefanus Kurniawan +1
Hallucinations are a key concern when creating applications that rely on Foundation models (FMs). Understanding where and how these subtle failures occur in an application relies o…
cs.CL2024
RAGProbe: An Automated Approach for Evaluating RAG Applications
Shangeetha Sivasothy, Scott Barnett, Stefanus Kurniawan +2
Retrieval Augmented Generation (RAG) is increasingly being used when building Generative AI applications. Evaluating these applications and RAG pipelines is mostly done manually, v…
cs.CL2024
Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models
Scott Barnett, Zac Brannelly, Stefanus Kurniawan +1
Large Language Models (LLMs) have the unique capability to understand and generate human-like text from input queries. When fine-tuned, these models show enhanced performance on do…