1 citations · 1 across the 1 of their papers we have counts for
8 papers
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemm…
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
Maxime Guigon, Lucas Dixon, Michaël E. Sander
Knowledge Distillation (KD) is a critical tool for training Large Language Models (LLMs), yet the majority of research focuses on approaches that rely solely on output logits, negl…
MIND: Monge Inception Distance for Generative Models Evaluation
Quentin Berthet, Yu-Han Wu, Clement Crepy +3
We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). Th…
Clustering in Deep Stochastic Transformers
Lev Fedorov, Michaël E. Sander, Romuald Elie +2
Transformers have revolutionized deep learning across various domains but understanding the precise token dynamics remains a theoretical challenge. Existing theories of deep Transf…
Differentiable Knapsack and Top-k Operators via Dynamic Programming
Germain Vivier-Ardisson, Michaël E. Sander, Axel Parmentier +1
Knapsack and Top-k operators are useful for selecting discrete subsets of variables. However, their integration into neural networks is challenging as they are piecewise constant,…
Joint Learning of Energy-based Models and their Partition Function
Michael E. Sander, Vincent Roulet, Tianlin Liu +1
Energy-based models (EBMs) offer a flexible framework for parameterizing probability distributions using neural networks. However, learning EBMs by exact maximum likelihood estimat…