8 papers
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
Dongyang Fan, Diba Hashemi, Sai Praneeth Karimireddy +1
Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prior work highlighted only one usef…
HalluHard: A Hard Multi-Turn Hallucination Benchmark
Dongyang Fan, Sebastien Delsad, Nicolas Flammarion +1
Large language models (LLMs) still produce plausible-sounding but ungrounded factual claims, a problem that worsens in multi-turn dialogue as context grows and early errors cascade…
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
Dongyang Fan, Vinko SabolÄec, Martin Jaggi
Large Language Models (LLMs) are commonly pretrained on vast corpora of text without utilizing contextual metadata such as source, quality, or topic, leading to a context-free lear…
Do Data Valuations Make Good Data Prices?
Dongyang Fan, Tyler J. Rotello, Sai Praneeth Karimireddy
As large language models increasingly rely on external data sources, compensating data contributors has become a central concern. But how should these payments be devised? We revis…
TiMoE: Time-Aware Mixture of Language Experts
Robin Faro, Dongyang Fan, Tamar Alphaidze +1
Large language models (LLMs) are typically trained on fixed snapshots of the web, which means that their knowledge becomes stale and their predictions risk temporal leakage: relyin…