6 papers · 1 filter
A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage
Francis Bach
We consider supervised learning problems in which set predictions provide explicit uncertainty estimates. Using Choquet integrals (a.k.a. Lov{á}sz extensions), we propose a convex…
On the Effectiveness of the z-Transform Method in Quadratic Optimization
Francis Bach
The z-transform of a sequence is a classical tool used within signal processing, control theory, computer science, and electrical engineering. It allows for studying sequences from…
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
Frederik Kunstner, Francis Bach
Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by…
Spectral structure learning for clinical time series
Ivan Lerner, Anita Burgun, Francis Bach
We develop and evaluate a structure learning algorithm for clinical time series. Clinical time series are multivariate time series observed in multiple patients and irregularly sam…
An Uncertainty Principle for Linear Recurrent Neural Networks
Alexandre François, Antonio Orvieto, Francis Bach
We consider linear recurrent neural networks, which have become a key building block of sequence modeling due to their ability for stable and effective long-range modeling. In this…
Optimizing Estimators of Squared Calibration Errors in Classification
Sebastian G. Gruber, Francis Bach
In this work, we propose a mean-squared error-based risk that enables the comparison and optimization of estimators of squared calibration errors in practical settings. Improving t…