7 papers
Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100
We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
Xinyu Zhou, Simin Fan, Martin Jaggi
Influence functions provide a principled method to assess the contribution of individual training samples to a specific target. Yet, their high computational costs limit their appl…
Semantic uncertainty in advanced decoding methods for LLM generation
Darius Foodeei, Simin Fan, Martin Jaggi
This study investigates semantic uncertainty in large language model (LLM) outputs across different decoding methods, focusing on emerging techniques like speculative sampling and…
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Shang-Chi Tsai, Yun-Nung Chen
With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
Simin Fan, Maria Ios Glarou, Martin Jaggi
The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existi…
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
Xinyu Zhou, Simin Fan, Martin Jaggi +1
Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGr…