papers

Publications (14)

cs.LG2026

The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning

Simin Fan, Dimitris Paparas, Natasha Noy +3

Understanding how language model capabilities transfer from pretraining to supervised fine-tuning (SFT) is fundamental to efficient model development and data curation. In this wor…

cs.LG2025

HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation

Xinyu Zhou, Simin Fan, Martin Jaggi

Influence functions provide a principled method to assess the contribution of individual training samples to a specific target. Yet, their high computational costs limit their appl…

cs.CY2024

Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants

Beatriz Borges, Negar Foroutan, Deniz Bayazit +87

AI assistants are being increasingly used by students enrolled in higher education institutions. While these tools provide opportunities for improved teaching and education, they a…

cs.CL2023

Irreducible Curriculum for Language Model Pretraining

Simin Fan, Martin Jaggi

Automatic data selection and curriculum design for training large language models is challenging, with only a few existing methods showing improvements over standard training. Furt…

cs.LG2025

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

Simin Fan, Maria Ios Glarou, Martin Jaggi

The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existi…

cs.CL2025

Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

David Grangier, Simin Fan, Skyler Seto +1

Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain…

cs.CL2025

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100

We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…

cs.CL2023

MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Zeming Chen, Alejandro Hernández Cano, Angelika Romanou +17

Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning…

cs.CL2025

Semantic uncertainty in advanced decoding methods for LLM generation

Darius Foodeei, Simin Fan, Martin Jaggi

This study investigates semantic uncertainty in large language model (LLM) outputs across different decoding methods, focusing on emerging techniques like speculative sampling and…

cs.LG2024

Dynamic Gradient Alignment for Online Data Mixing

Simin Fan, David Grangier, Pierre Ablin

The composition of training data mixtures is critical for effectively training large language models (LLMs), as it directly impacts their performance on downstream tasks. Our goal…

cs.LG2024

DoGE: Domain Reweighting with Generalization Estimation

Simin Fan, Matteo Pagliardini, Martin Jaggi

The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rel…

cs.LG2024

Deep Grokking: Would Deep Neural Networks Generalize Better?

Simin Fan, Razvan Pascanu, Martin Jaggi

Recent research on the grokking phenomenon has illuminated the intricacies of neural networks' training dynamics and their generalization behaviors. Grokking refers to a sharp rise…

cs.HC2022

Towards Process-Oriented, Modular, and Versatile Question Generation that Meets Educational Needs

Xu Wang, Simin Fan, Jessica Houghton +1

NLP-powered automatic question generation (QG) techniques carry great pedagogical potential of saving educators' time and benefiting student learning. Yet, QG systems have not been…

cs.LG2025

NeuralGrok: Accelerate Grokking by Neural Gradient Transformation

Xinyu Zhou, Simin Fan, Martin Jaggi +1

Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGr…