30 citations · 111 across the 48 of their papers we have counts for
47 papers · 1 filter
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Yaswanth Chittepu, Ativ Joshi, Sohini Chintala +1
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-ti…
The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2
Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…
Training ML Models with Predictable Failures
Will Schwarzer, Scott Niekum
Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely large enough to observe the f…
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee +1
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost…
Adaptive Margin RLHF via Preference over Preferences
Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Yaswanth Chittepu, Blossom Metevier, Will Schwarzer +3
Existing approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains. To ensure relia…