4 papers · 1 filter
DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain
Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama +2
We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents span…
GaelEval: Benchmarking LLM Performance for Scottish Gaelic
Peter Devine, William Lamb, Beatrice Alex +5
Multilingual large language models (LLMs) often exhibit emergent 'shadow' capabilities in languages without official support, yet their performance on these languages remains uneve…
Kakugo: Distillation of Low-Resource Languages into Small Language Models
Peter Devine, Mardhiyah Sanni, Farid Adilazuarda +2
We present Kakugo, a novel and cost-effective pipeline designed to train general-purpose Small Language Models (SLMs) for low-resource languages using only the language name as inp…
M-IFEval: Multilingual Instruction-Following Evaluation
Antoine Dussolle, Andrea Cardeña DÃaz, Shota Sato +1
Instruction following is a core capability of modern Large language models (LLMs), making evaluating this capability essential to understanding these models. The Instruction Follow…