16 papers
CalArena: A Large-Scale Post-Hoc Calibration Benchmark
Eugène Berta, David Holzmüller, Francis Bach +1
Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated. Post-hoc calibration provides a simple and wi…
Conditional Coverage Diagnostics for Conformal Prediction
Sacha Braun, David Holzmüller, Michael I. Jordan +1
Evaluating conditional coverage remains one of the most persistent challenges in assessing the reliability of predictive systems. Although conformal methods can give guarantees on…
STRABLE: Benchmarking Tabular Machine Learning with Strings
Gioia Blayer, Myung Jun Kim, Félix Lefebvre +8
Benchmarking tabular learning has revealed the benefit of dedicated architectures, pushing the state of the art. But real-world tables often contain string entries, beyond numbers,…
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
Alan Arazi, Eilam Shapira, Shoham Grunblat +8
Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numeric…
Beyond ReLU: How Activations Affect Neural Kernels and Random Wide Networks
David Holzmüller, Max Schölpple
In recent years, the neural tangent kernel (NTK) and neural network Gaussian process kernel (NNGP) have given theoreticians tractable limiting cases of fully connected neural netwo…
Convergence Rates for Non-Log-Concave Sampling and Log-Partition Estimation
David Holzmüller, Francis Bach
Sampling from Gibbs distributions and computing their log-partition function are fundamental tasks in statistics, machine learning, and statistical physics. While efficient algorit…