activity
20242026
most citedVerify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.LG2024

Large Language Models can be Strong Self-Detoxifiers

Ching-Yun Ko, Pin-Yu Chen, Payel Das +6

Reducing the likelihood of generating harmful and toxic output is an essential task when aligning large language models (LLMs). Existing methods mainly rely on training an external…

stat.ML2024

Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking

Gabriel Rioux, Apoorva Nitsure, Mattia Rigotti +2

Stochastic dominance is an important concept in probability theory, econometrics and social choice theory for robustly modeling agents' preferences between random outcomes. While m…

cs.LG2024

Information Theoretic Guarantees For Policy Alignment In Large Language Models

Youssef Mroueh

Policy alignment of large language models refers to constrained policy optimization, where the policy is optimized to maximize a reward while staying close to a reference policy wi…

cs.LG2024

Distributional Preference Alignment of LLMs via Optimal Transport

Igor Melnyk, Youssef Mroueh, Brian Belgodere +6

Current LLM alignment techniques use pairwise human preferences at a sample level, and as such, they do not imply an alignment on the distributional level. We propose in this paper…

cs.LG2024

Risk Aware Benchmarking of Large Language Models

Apoorva Nitsure, Youssef Mroueh, Mattia Rigotti +6

We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical…

cs.LG2024

Auditing and Generating Synthetic Data with Controllable Trust Trade-offs

Brian Belgodere, Pierre Dognin, Adam Ivankay +11

Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate…