5 papers
Reliability Scaling Laws for Quantized Large Language Models
Sirine Ayadi, Sándor Daróczi, Stephan Günnemann +1
Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve st…
Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation
Johanna Sommer, John Rachwan, Nils Fleischmann +2
Flow matching models generate high-fidelity molecular geometries but incur significant computational costs during inference, requiring hundreds of network evaluations. This inferen…
Uncertainty for Active Learning on Graphs
Dominik Fuchsgruber, Tom Wollschläger, Bertrand Charpentier +2
Uncertainty Sampling is an Active Learning strategy that aims to improve the data efficiency of machine learning models by iteratively acquiring labels of data points with the high…
Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks Using the Marginal Likelihood
Rayen Dhahri, Alexander Immer, Betrand Charpentier +2
Neural network sparsification is a promising avenue to save computational time and memory costs, especially in an age where many successful AI models are becoming too large to naï…
Predicting Probabilities of Error to Combine Quantization and Early Exiting: QuEE
Florence Regol, Joud Chataoui, Bertrand Charpentier +3
Machine learning models can solve complex tasks but often require significant computational resources during inference. This has led to the development of various post-training com…