3 papers
cs.LG2026
Selective Safety Steering via Value-Filtered Decoding
Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh +2
While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by…
cs.LG2025
Semi-Supervised Risk Control via Prediction-Powered Inference
Bat-Sheva Einbinder, Liran Ringel, Yaniv Romano
The risk-controlling prediction sets (RCPS) framework is a general tool for transforming the output of any machine learning model to design a predictive rule with rigorous error ra…
cs.LG2024
Label Noise Robustness of Conformal Prediction
Bat-Sheva Einbinder, Shai Feldman, Stephen Bates +3
We study the robustness of conformal prediction, a powerful tool for uncertainty quantification, to label noise. Our analysis tackles both regression and classification problems, c…