123 citations · 123 across the 4 of their papers we have counts for
5 papers · 1 filter
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
Simin Fan, Maria Ios Glarou, Martin Jaggi
The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existi…
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
Xinyu Zhou, Simin Fan, Martin Jaggi +1
Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGr…
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
Xinyu Zhou, Simin Fan, Martin Jaggi
Influence functions provide a principled method to assess the contribution of individual training samples to a specific target. Yet, their high computational costs limit their appl…
Deep Grokking: Would Deep Neural Networks Generalize Better?
Simin Fan, Razvan Pascanu, Martin Jaggi
Recent research on the grokking phenomenon has illuminated the intricacies of neural networks' training dynamics and their generalization behaviors. Grokking refers to a sharp rise…
DoGE: Domain Reweighting with Generalization Estimation
Simin Fan, Matteo Pagliardini, Martin Jaggi
The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rel…