collaborators

5 papers

cs.CL2026

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama +2

We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents span…

cs.CL2026

GaelEval: Benchmarking LLM Performance for Scottish Gaelic

Peter Devine, William Lamb, Beatrice Alex +5

Multilingual large language models (LLMs) often exhibit emergent 'shadow' capabilities in languages without official support, yet their performance on these languages remains uneve…

cs.CL2026

Kakugo: Distillation of Low-Resource Languages into Small Language Models

Peter Devine, Mardhiyah Sanni, Farid Adilazuarda +2

We present Kakugo, a novel and cost-effective pipeline designed to train general-purpose Small Language Models (SLMs) for low-resource languages using only the language name as inp…

cs.CL2025

M-IFEval: Multilingual Instruction-Following Evaluation

Antoine Dussolle, Andrea Cardeña Díaz, Shota Sato +1

Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Follow…

cs.LG2025

ALoFTRAG: Automatic Local Fine Tuning for Retrieval Augmented Generation

Peter Devine

Retrieval Augmented Generation (RAG) systems have been shown to improve the accuracy of Large Language Model (LLM) outputs. However, these models can often achieve low accuracy whe…