Publications (48)
Sinkformers: Transformers with Doubly Stochastic Attention
Michael E. Sander, Pierre Ablin, Mathieu Blondel +1
Learning with Differentiable Perturbed Optimizers
Quentin Berthet, Mathieu Blondel, Olivier Teboul +3
Blind Source Separation with Optimal Transport Non-negative Matrix Factorization
Antoine Rolet, Vivien Seguy, Mathieu Blondel +1
Differentiable Knapsack and Top-k Operators via Dynamic Programming
Germain Vivier-Ardisson, Michaël E. Sander, Axel Parmentier +1
Efficient and Modular Implicit Differentiation
Mathieu Blondel, Quentin Berthet, Marco Cuturi +5
Large-Scale Optimal Transport and Mapping Estimation
Vivien Seguy, Bharath Bhushan Damodaran, Rémi Flamary +3
Learning with Fitzpatrick Losses
Seta Rakotomandimby, Jean-Philippe Chancelier, Michel de Lara +1
Structured Prediction with Projection Oracles
Mathieu Blondel
How do Transformers perform In-Context Autoregressive Learning?
Michael E. Sander, Raja Giryes, Taiji Suzuki +2
Fast Differentiable Sorting and Ranking
Mathieu Blondel, Olivier Teboul, Quentin Berthet +1
Soft-DTW: a Differentiable Loss Function for Time-Series
Marco Cuturi, Mathieu Blondel
On Teacher Hacking in Language Model Distillation
Daniil Tiapkin, Daniele Calandriello, Johan Ferret +4
Learning with Local Search MCMC Layers
Germain Vivier-Ardisson, Mathieu Blondel, Axel Parmentier
Regularized Large Neighborhood Search
Germain Vivier-Ardisson, Laurent Demonet, Axel Parmentier +1
Polynomial Networks and Factorization Machines: New Insights and Efficient Training Algorithms
Mathieu Blondel, Masakazu Ishihata, Akinori Fujino +1
Routers in Vision Mixture of Experts: An Empirical Study
Tianlin Liu, Mathieu Blondel, Carlos Riquelme +1
Loss Functions and Operators Generated by f-Divergences
Vincent Roulet, Tianlin Liu, Nino Vieillard +2
Learning with Fenchel-Young Losses
Mathieu Blondel, André F. T. Martins, Vlad Niculae
Stepping on the Edge: Curvature Aware Learning Rate Tuners
Vincent Roulet, Atish Agarwala, Jean-Bastien Grill +3
A Regularized Framework for Sparse and Structured Neural Attention
Vlad Niculae, Mathieu Blondel
Self-Supervised Learning of Audio Representations from Permutations with Differentiable Ranking
Andrew N Carr, Quentin Berthet, Mathieu Blondel +2
Differentiable Dynamic Programming for Structured Prediction and Attention
Arthur Mensch, Mathieu Blondel
Higher-Order Factorization Machines
Mathieu Blondel, Akinori Fujino, Naonori Ueda +1
Learning Energy Networks with Generalized Fenchel-Young Losses
Mathieu Blondel, Felipe Llinares-López, Robert Dadashi +2
Sparse Continuous Distributions and Fenchel-Young Losses
André F. T. Martins, Marcos Treviso, António Farinhas +4
Smooth and Sparse Optimal Transport
Mathieu Blondel, Vivien Seguy, Antoine Rolet
Momentum Residual Neural Networks
Michael E. Sander, Pierre Ablin, Mathieu Blondel +1
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
The Elements of Differentiable Programming
Mathieu Blondel, Vincent Roulet
Differentiable Divergences Between Time Series
Mathieu Blondel, Arthur Mensch, Jean-Philippe Vert
Sparsity-Constrained Optimal Transport
Tianlin Liu, Joan Puigcerver, Mathieu Blondel
Cutting Some Slack for SGD with Adaptive Polyak Stepsizes
Robert M. Gower, Mathieu Blondel, Nidham Gazagnadou +1
Multi-output Polynomial Networks and Factorization Machines
Mathieu Blondel, Vlad Niculae, Takuma Otsuka +1
Implicit Diffusion: Efficient Optimization through Stochastic Sampling
Pierre Marion, Anna Korba, Peter Bartlett +6
Decoding-time Realignment of Language Models
Tianlin Liu, Shangmin Guo, Leonardo Bianco +7
Dual Gauss-Newton Directions for Deep Learning
Vincent Roulet, Mathieu Blondel
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
Geometric Losses for Distributional Learning
Arthur Mensch, Mathieu Blondel, Gabriel Peyré
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson +2
Implicit differentiation of Lasso-type models for hyperparameter optimization
Quentin Bertrand, Quentin Klopfenstein, Mathieu Blondel +3
Learning Classifiers with Fenchel-Young Losses: Generalized Entropies, Margins, and Algorithms
Mathieu Blondel, André F. T. Martins, Vlad Niculae
Fast, Differentiable and Sparse Top-k: a Convex Analysis Perspective
Michael E. Sander, Joan Puigcerver, Josip Djolonga +2
SparseMAP: Differentiable Sparse Structured Inference
Vlad Niculae, André F. T. Martins, Mathieu Blondel +1
Scikit-learn: Machine Learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort +16
API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel +12
Direct Language Model Alignment from Online AI Feedback
Shangmin Guo, Biao Zhang, Tianlin Liu +9
Joint Learning of Energy-based Models and their Partition Function
Michael E. Sander, Vincent Roulet, Tianlin Liu +1
Implicit differentiation for fast hyperparameter selection in non-smooth convex learning
Quentin Bertrand, Quentin Klopfenstein, Mathurin Massias +4