6 papers
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
Thomas F Burns, Letitia Parcalabescu, Stephan Wäldchen +5
Scaling data quantity is essential for large language models (LLMs), yet recent findings show that data quality can significantly boost performance and training efficiency. We intr…
A Family of LLMs Liberated from Static Vocabularies
Aleph Alpha, :, Adnen Abdessaied +35
Tokenization is a central component of natural language processing in current large language models (LLMs), enabling models to convert raw text into processable units. Although lea…
Measuring Language Model Hallucinations Through Distributional Correctness
Thomas F Burns
Common evaluation paradigms for language models focus on scoring single responses through accuracy metrics or proper scoring rules, failing to capture the full richness of a model'…
Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
Thomas F Burns, Tomoki Fukai, Christopher J Earls
Large language models (LLMs) demonstrate an impressive ability to utilise information within the context of their input sequences to appropriately respond to data unseen by the LLM…
Generative AI Policy and Governance Considerations for Health Security in Southeast Asia
Thomas F Burns
Southeast Asia is a geopolitically and socio-economically significant region with unique challenges and opportunities. Intensifying progress in generative AI against a backdrop of…
Semantically-correlated memories in a dense associative model
Thomas F Burns
I introduce a novel associative memory model named Correlated Dense Associative Memory (CDAM), which integrates both auto- and hetero-association in a unified framework for continu…