1 citations · 1 across the 3 of their papers we have counts for
3 papers
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
Syrine Belakaria, Joshua Kazdan, Charles Marx +5
Reinforcement learning from human feedback (RLHF) has become a cornerstone of the training and alignment pipeline for large language models (LLMs). Recent advances, such as direct…
Calibration by Distribution Matching: Trainable Kernel Calibration Metrics
Charles Marx, Sofian Zalouk, Stefano Ermon
Calibration ensures that probabilistic forecasts meaningfully capture uncertainty by requiring that predicted probabilities align with empirical frequencies. However, many existing…
Modular Conformal Calibration
Charles Marx, Shengjia Zhao, Willie Neiswanger +1
Uncertainty estimates must be calibrated (i.e., accurate) and sharp (i.e., informative) in order to be useful. This has motivated a variety of methods for recalibration, which use…