activity
20232025
most citedMEDITRON-70B: Scaling Medical Pretraining for Large Language Models

123 citations · 123 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

Simin Fan, Maria Ios Glarou, Martin Jaggi

The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existi…

cs.LG2025

NeuralGrok: Accelerate Grokking by Neural Gradient Transformation

Xinyu Zhou, Simin Fan, Martin Jaggi +1

Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGr…

cs.LG2024

HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation

Xinyu Zhou, Simin Fan, Martin Jaggi

Influence functions provide a principled method to assess the contribution of individual training samples to a specific target. Yet, their high computational costs limit their appl…

cs.LG2024

Deep Grokking: Would Deep Neural Networks Generalize Better?

Simin Fan, Razvan Pascanu, Martin Jaggi

Recent research on the grokking phenomenon has illuminated the intricacies of neural networks' training dynamics and their generalization behaviors. Grokking refers to a sharp rise…

cs.LG2023

DoGE: Domain Reweighting with Generalization Estimation

Simin Fan, Matteo Pagliardini, Martin Jaggi

The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs). Despite its importance, recent LLMs still rel…