Publications (71)
Implicit Kernel Learning
Chun-Liang Li, Wei-Cheng Chang, Youssef Mroueh +2
Kernels are powerful and versatile tools in machine learning and statistics. Although the notion of universal kernels and characteristic kernels has been studied, kernel selection…
Learning Implicit Text Generation via Feature Matching
Inkit Padhi, Pierre Dognin, Ke Bai +4
Generative feature matching network (GFMN) is an approach for training implicit generative models for images by performing moment matching on features from pre-trained neural netwo…
Auditing Differential Privacy in High Dimensions with the Kernel Quantum Rényi Divergence
Carles Domingo-Enrich, Youssef Mroueh
Differential privacy (DP) is the de facto standard for private data release and private machine learning. Auditing black-box DP algorithms and mechanisms to certify whether they sa…
Sobolev Descent
Youssef Mroueh, Tom Sercu, Anant Raj
We study a simplification of GAN training: the problem of transporting particles from a source to a target distribution. Starting from the Sobolev GAN critic, part of the gradient…
Sobolev Independence Criterion
Youssef Mroueh, Tom Sercu, Mattia Rigotti +2
We propose the Sobolev Independence Criterion (SIC), an interpretable dependency measure between a high dimensional random variable X and a response variable Y . SIC decomposes to…
Multiclass Learning with Simplex Coding
Youssef Mroueh, Tomaso Poggio, Lorenzo Rosasco +1
In this paper we discuss a novel framework for multiclass learning, defined by a suitable coding/decoding strategy, namely the simplex coding, that allows to generalize to multiple…
GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance
Jiri Navratil, Jarret Ross, Payel Das +4
The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology…
Learning with Group Invariant Features: A Kernel Perspective
Youssef Mroueh, Stephen Voinea, Tomaso Poggio
We analyze in this paper a random feature map based on a theory of invariance I-theory introduced recently. More specifically, a group invariant signal signature is obtained throug…
Effective Dynamics of Generative Adversarial Networks
Steven Durr, Youssef Mroueh, Yuhai Tu +1
Generative adversarial networks (GANs) are a class of machine-learning models that use adversarial training to generate new samples with the same (potentially very complex) statist…
Physics-enhanced deep surrogates for partial differential equations
Raphaël Pestourie, Youssef Mroueh, Chris Rackauckas +2
Many physics and engineering applications demand Partial Differential Equations (PDE) property evaluations that are traditionally computed with resource-intensive high-fidelity num…
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
Gholamali Aminian, Idan Shenfeld, Amir R. Asadi +2
A simple yet effective method for inference-time alignment of generative models is Best-of- (BoN), where outcomes are sampled from a reference policy, evaluated using a prox…
Optimizing Functionals on the Space of Probabilities with Input Convex Neural Networks
David Alvarez-Melis, Yair Schiff, Youssef Mroueh
Gradient flows are a powerful tool for optimizing functionals in general metric spaces, including the space of probabilities endowed with the Wasserstein metric. A typical approach…
Fisher GAN
Youssef Mroueh, Tom Sercu
Generative Adversarial Networks (GANs) are powerful models for learning complex distributions. Stable training of GANs has been addressed in many recent works which explore differe…
McGan: Mean and Covariance Feature Matching GAN
Youssef Mroueh, Tom Sercu, Vaibhava Goel
We introduce new families of Integral Probability Metrics (IPM) for training Generative Adversarial Networks (GAN). Our IPMs are based on matching statistics of distributions embed…
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis
We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of- test-time scaling with…
Learning Implicit Generative Models by Matching Perceptual Features
Cicero Nogueira dos Santos, Youssef Mroueh, Inkit Padhi +1
Perceptual features (PFs) have been used with great success in tasks such as transfer learning, style transfer, and super-resolution. However, the efficacy of PFs as key source of…
On the Convergence of Gradient Descent in GANs: MMD GAN As a Gradient Flow
Youssef Mroueh, Truyen Nguyen
We consider the maximum mean discrepancy () GAN problem and propose a parametric kernelized gradient flow that mimics the min-max game in gradient regularized $\mathr…
Local Group Invariant Representations via Orbit Embeddings
Anant Raj, Abhishek Kumar, Youssef Mroueh +2
Invariance to nuisance transformations is one of the desirable properties of effective representations. We consider transformations that form a \emph{group} and propose an approach…
Deep Multimodal Learning for Audio-Visual Speech Recognition
Youssef Mroueh, Etienne Marcheret, Vaibhava Goel
In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an app…
q-ary Compressive Sensing
Youssef Mroueh, Lorenzo Rosasco
We introduce q-ary compressive sensing, an extension of 1-bit compressive sensing. We propose a novel sensing mechanism and a corresponding recovery procedure. The recovery propert…
Cloud-Based Real-Time Molecular Screening Platform with MolFormer
Brian Belgodere, Vijil Chenthamarakshan, Payel Das +9
With the prospect of automating a number of chemical tasks with high fidelity, chemical language processing models are emerging at a rapid speed. Here, we present a cloud-based rea…
Alleviating Noisy Data in Image Captioning with Cooperative Distillation
Pierre Dognin, Igor Melnyk, Youssef Mroueh +4
Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their…
Co-Occuring Directions Sketching for Approximate Matrix Multiply
Youssef Mroueh, Etienne Marcheret, Vaibhava Goel
We introduce co-occurring directions sketching, a deterministic algorithm for approximate matrix product (AMM), in the streaming model. We show that co-occuring directions achieves…
Measuring Generalization with Optimal Transport
Ching-Yao Chuang, Youssef Mroueh, Kristjan Greenewald +2
Understanding the generalization of deep neural networks is one of the most important tasks in deep learning. Although much progress has been made, theoretical error bounds still o…
Convex Learning of Multiple Tasks and their Structure
Carlo Ciliberto, Youssef Mroueh, Tomaso Poggio +1
Reducing the amount of human supervision is a key problem in machine learning and a natural approach is that of exploiting the relations (structure) among different tasks. This is…
Gromov-Wasserstein Distances: Entropic Regularization, Duality, and Sample Complexity
Zhengxin Zhang, Ziv Goldfeld, Youssef Mroueh +1
The Gromov-Wasserstein (GW) distance, rooted in optimal transport (OT) theory, quantifies dissimilarity between metric measure spaces and provides a framework for aligning heteroge…
Robust Phase Retrieval and Super-Resolution from One Bit Coded Diffraction Patterns
Youssef Mroueh
In this paper we study a realistic setup for phase retrieval, where the signal of interest is modulated or masked and then for each modulation or mask a diffraction pattern is coll…
Random Maxout Features
Youssef Mroueh, Steven Rennie, Vaibhava Goel
In this paper, we propose and study random maxout features, which are constructed by first projecting the input data onto sets of randomly generated vectors with Gaussian elements,…
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Youssef Mroueh, Nicolas Dupuis, Brian Belgodere +6
We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Opti…
Can a biologically-plausible hierarchy effectively replace face detection, alignment, and recognition pipelines?
Qianli Liao, Joel Z Leibo, Youssef Mroueh +1
The standard approach to unconstrained face recognition in natural photographs is via a detection, alignment, recognition pipeline. While that approach has achieved impressive resu…
Unsupervised Hierarchy Matching with Optimal Transport over Hyperbolic Spaces
David Alvarez-Melis, Youssef Mroueh, Tommi S. Jaakkola
This paper focuses on the problem of unsupervised alignment of hierarchical data such as ontologies or lexical databases. This is a problem that appears across areas, from natural…
Sobolev GAN
Youssef Mroueh, Chun-Liang Li, Tom Sercu +2
We propose a new Integral Probability Metric (IPM) between distributions: the Sobolev IPM. The Sobolev IPM compares the mean discrepancy of two distributions for functions (critic)…
Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization
Clément Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh
Optimizing functionals over the space of probability measures is now ubiquitous in machine learning. A widely used approach is to perform the optimization directly over the Wassers…
Learning with Stochastic Orders
Carles Domingo-Enrich, Yair Schiff, Youssef Mroueh
Learning high-dimensional distributions is often done with explicit likelihood modeling or implicit modeling via minimizing integral probability metrics (IPMs). In this paper, we e…
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
Youssef Mroueh, Carlos Fonseca, Brian Belgodere +1
Scientific algorithm discovery is iterative: hypotheses are proposed, implemented, stress-tested, and revised. Current LLM-guided search systems accelerate proposal generation, but…
Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
Nicolas Dupuis, Adarsh Tiwari, Youssef Mroueh +3
Qiskit is an open-source quantum computing framework that allows users to design, simulate, and run quantum circuits on real quantum hardware. We explore post-training techniques f…
Information Theoretic Guarantees For Policy Alignment In Large Language Models
Youssef Mroueh
Policy alignment of large language models refers to constrained policy optimization, where the policy is optimized to maximize a reward while staying close to a reference policy wi…
Generative Modeling with Denoising Auto-Encoders and Langevin Sampling
Adam Block, Youssef Mroueh, Alexander Rakhlin
We study convergence of a generative modeling method that first estimates the score function of the distribution using Denoising Auto-Encoders (DAE) or Denoising Score Matching (DS…
Asymmetrically Weighted CCA And Hierarchical Kernel Sentence Embedding For Image & Text Retrieval
Youssef Mroueh, Etienne Marcheret, Vaibhava Goel
Joint modeling of language and vision has been drawing increasing interest. A multimodal data representation allowing for bidirectional retrieval of images by sentences and vice ve…
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
Youssef Mroueh
Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calib…
Large-Scale Chemical Language Representations Capture Molecular Structure and Properties
Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan +3
Models based on machine learning can enable accurate and fast molecular property predictions, which is of interest in drug discovery and material design. Various supervised machine…
Gradient Flows and Riemannian Structure in the Gromov-Wasserstein Geometry
Zhengxin Zhang, Ziv Goldfeld, Kristjan Greenewald +2
The Wasserstein space of probability measures is known for its intricate Riemannian structure, which underpins the Wasserstein geometry and enables gradient flow algorithms. Howeve…
Quantization and Greed are Good: One bit Phase Retrieval, Robustness and Greedy Refinements
Youssef Mroueh, Lorenzo Rosasco
In this paper, we study the problem of robust phase recovery. We investigate a novel approach based on extremely quantized (one-bit) phase-less measurements and a corresponding rec…
Large Language Models can be Strong Self-Detoxifiers
Ching-Yun Ko, Pin-Yu Chen, Payel Das +6
Reducing the likelihood of generating harmful and toxic output is an essential task when aligning large language models (LLMs). Existing methods mainly rely on training an external…
Kernel Stein Generative Modeling
Wei-Cheng Chang, Chun-Liang Li, Youssef Mroueh +1
We are interested in gradient-based Explicit Generative Modeling where samples can be derived from iterative gradient updates based on an estimate of the score function of the data…
Semi-Supervised Learning with IPM-based GANs: an Empirical Study
Tom Sercu, Youssef Mroueh
We present an empirical investigation of a recent class of Generative Adversarial Networks (GANs) using Integral Probability Metrics (IPM) and their performance for semi-supervised…
A Decentralized Parallel Algorithm for Training Generative Adversarial Nets
Mingrui Liu, Wei Zhang, Youssef Mroueh +4
Generative Adversarial Networks (GANs) are a powerful class of generative models in the deep learning community. Current practice on large-scale GAN training utilizes large models…
Improving Efficiency in Large-Scale Decentralized Distributed Training
Wei Zhang, Xiaodong Cui, Abdullah Kayi +9
Decentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to p…
Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets
Mingrui Liu, Youssef Mroueh, Jerret Ross +4
Adaptive gradient algorithms perform gradient-based updates using the history of gradients and are ubiquitous in training deep neural networks. While adaptive gradient methods theo…
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
Khiem Le, Phuc Nguyen, Youssef Mroueh +4
Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critic…
Self-critical Sequence Training for Image Captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh +2
Recently it has been shown that policy-gradient methods for reinforcement learning can be utilized to train deep end-to-end systems directly on non-differentiable metrics for the t…
Fair Mixup: Fairness via Interpolation
Ching-Yao Chuang, Youssef Mroueh
Training classifiers under fairness constraints such as group fairness, regularizes the disparities of predictions between the groups. Nevertheless, even though the constraints are…
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
Yihao Xue, Kristjan Greenewald, Youssef Mroueh +1
Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-consistency-based techniques have been pro…
Tabular Transformers for Modeling Multivariate Time Series
Inkit Padhi, Yair Schiff, Igor Melnyk +6
Tabular datasets are ubiquitous in data science applications. Given their importance, it seems natural to apply state-of-the-art deep learning algorithms in order to fully unlock t…
Regularized Finite Dimensional Kernel Sobolev Discrepancy
Youssef Mroueh
We show in this note that the Sobolev Discrepancy introduced in Mroueh et al in the context of generative adversarial networks, is actually the weighted negative Sobolev norm $||.|…
Unbalanced Sobolev Descent
Youssef Mroueh, Mattia Rigotti
We introduce Unbalanced Sobolev Descent (USD), a particle descent algorithm for transporting a high dimensional source distribution to a target distribution that does not necessari…
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
Gholamali Aminian, Amir R. Asadi, Idan Shenfeld +1
Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underuti…
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
Gabriel Rioux, Apoorva Nitsure, Mattia Rigotti +2
Stochastic dominance is an important concept in probability theory, econometrics and social choice theory for robustly modeling agents' preferences between random outcomes. While m…
Wasserstein Barycenter Model Ensembling
Pierre Dognin, Igor Melnyk, Youssef Mroueh +3
In this paper we propose to perform model ensembling in a multiclass or a multilabel learning setting using Wasserstein (W.) barycenters. Optimal transport metrics, such as the Was…
Tighter Sparse Approximation Bounds for ReLU Neural Networks
Carles Domingo-Enrich, Youssef Mroueh
A well-known line of work (Barron, 1993; Breiman, 1993; Klusowski & Barron, 2018) provides bounds on the width of a ReLU two-layer neural network needed to approximate a functi…
Active learning of deep surrogates for PDEs: Application to metasurface design
Raphaël Pestourie, Youssef Mroueh, Thanh V. Nguyen +2
Surrogate models for partial-differential equations are widely used in the design of meta-materials to rapidly evaluate the behavior of composable components. However, the training…
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
Pierre Dognin, Igor Melnyk, Youssef Mroueh +6
Image captioning has recently demonstrated impressive progress largely owing to the introduction of neural network algorithms trained on curated dataset like MS-COCO. Often work in…
Adversarial Semantic Alignment for Improved Image Captions
Pierre L. Dognin, Igor Melnyk, Youssef Mroueh +2
In this paper we study image captioning as a conditional GAN training, proposing both a context-aware LSTM captioner and co-attentive discriminator, which enforces semantic alignme…
GP-MoLFormer: A Foundation Model For Molecular Generation
Jerret Ross, Brian Belgodere, Samuel C. Hoffman +4
Transformer-based models trained on large and general purpose datasets consisting of molecular strings have recently emerged as a powerful tool for successfully modeling various st…
Distributional Preference Alignment of LLMs via Optimal Transport
Igor Melnyk, Youssef Mroueh, Brian Belgodere +6
Current LLM alignment techniques use pairwise human preferences at a sample level, and as such, they do not imply an alignment on the distributional level. We propose in this paper…
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
Brian Belgodere, Pierre Dognin, Adam Ivankay +11
Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate…
Cycle Consistent Probability Divergences Across Different Spaces
Zhengxin Zhang, Youssef Mroueh, Ziv Goldfeld +1
Discrepancy measures between probability distributions are at the core of statistical inference and machine learning. In many applications, distributions of interest are supported…
Risk Aware Benchmarking of Large Language Models
Apoorva Nitsure, Youssef Mroueh, Mattia Rigotti +6
We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical…
Separation Results between Fixed-Kernel and Feature-Learning Probability Metrics
Carles Domingo-Enrich, Youssef Mroueh
Several works in implicit and explicit generative modeling empirically observed that feature-learning discriminators outperform fixed-kernel discriminators in terms of the sample q…
Fast Mixing of Multi-Scale Langevin Dynamics under the Manifold Hypothesis
Adam Block, Youssef Mroueh, Alexander Rakhlin +1
Recently, the task of image generation has attracted much attention. In particular, the recent empirical successes of the Markov Chain Monte Carlo (MCMC) technique of Langevin Dyna…
Wasserstein Style Transfer
Youssef Mroueh
We propose Gaussian optimal transport for Image style transfer in an Encoder/Decoder framework. Optimal transport for Gaussian measures has closed forms Monge mappings from source…