147 citations · 193 across the 32 of their papers we have counts for
22 papers · 1 filter
An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
Roberto Campbell, Momin Abbass, Muneeza Azmat +5
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. E…
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3
The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and test…
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
Rachel Ma, Dylan Hadfield-Menell, Kristjan Greenewald
Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the fir…
Differentially Private Wasserstein Barycenters
Anming Gu, Sasidhar Kunapuli, Mark Bun +2
The Wasserstein barycenter is defined as the mean of a set of probability measures under the optimal transport metric, and has numerous applications spanning machine learning, stat…
Entropic Causal Inference: Graph Identifiability
Spencer Compton, Kristjan Greenewald, Dmitriy Katz +1
Entropic causal inference is a recent framework for learning the causal graph between two variables from observational data by finding the information-theoretically simplest struct…
Private Continuous-Time Synthetic Trajectory Generation via Mean-Field Langevin Dynamics
Anming Gu, Edward Chien, Kristjan Greenewald
We provide an algorithm to privately generate continuous-time data (e.g. marginals from stochastic differential equations), which has applications in highly sensitive domains invol…