collaborators

8 papers

cs.CL2026

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

Dongyang Fan, Diba Hashemi, Sai Praneeth Karimireddy +1

Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prior work highlighted only one usef…

cs.AI2026

HalluHard: A Hard Multi-Turn Hallucination Benchmark

Dongyang Fan, Sebastien Delsad, Nicolas Flammarion +1

Large language models (LLMs) still produce plausible-sounding but ungrounded factual claims, a problem that worsens in multi-turn dialogue as context grows and early errors cascade…

cs.CL2025

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100

We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…

cs.CL2025

URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training

Dongyang Fan, Vinko Sabolčec, Martin Jaggi

Large Language Models (LLMs) are commonly pretrained on vast corpora of text without utilizing contextual metadata such as source, quality, or topic, leading to a context-free lear…

cs.GT2025

Do Data Valuations Make Good Data Prices?

Dongyang Fan, Tyler J. Rotello, Sai Praneeth Karimireddy

As large language models increasingly rely on external data sources, compensating data contributors has become a central concern. But how should these payments be devised? We revis…

cs.CL2025

TiMoE: Time-Aware Mixture of Language Experts

Robin Faro, Dongyang Fan, Tamar Alphaidze +1

Large language models (LLMs) are typically trained on fixed snapshots of the web, which means that their knowledge becomes stale and their predictions risk temporal leakage: relyin…