5 papers
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
Roland Daynauth, Christopher Clarke, Krisztian Flautner +2
The LLM-as-a-Judge paradigm offers a scalable, reference-free approach for evaluating language models. Although several calibration techniques have been proposed to better align th…
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
Roland Daynauth, Christopher Clarke, Krisztian Flautner +2
Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entai…
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
Roland Daynauth, Jason Mars
The SLAM paper demonstrated that on-device Small Language Models (SLMs) are a viable and cost-effective alternative to API-based Large Language Models (LLMs), such as OpenAI's GPT-…
Guylingo: The Republic of Guyana Creole Corpora
Christopher Clarke, Roland Daynauth, Charlene Wilkinson +2
While major languages often enjoy substantial attention and resources, the linguistic diversity across the globe encompasses a multitude of smaller, indigenous, and regional langua…
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
Chandra Irugalbandara, Ashish Mahendra, Roland Daynauth +6
Many companies use large language models (LLMs) offered as a service, like OpenAI's GPT-4, to create AI-enabled product experiences. Along with the benefits of ease-of-use and shor…