147 citations · 219 across the 49 of their papers we have counts for
3 papers · 2 filters
An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
Roberto Campbell, Momin Abbas, Momin Abbass +6
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. E…
Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3
The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and test…
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
Rachel Ma, Dylan Hadfield-Menell, Kristjan Greenewald
Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the fir…