activity
20162026
most citedExtrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

30 citations · 111 across the 48 of their papers we have counts for

collaborators
Showing cs.LGShow all

47 papers · 1 filter

cs.LG2026

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

Yaswanth Chittepu, Ativ Joshi, Sohini Chintala +1

Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-ti…

cs.LG2026

The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Sankaran Vaidyanathan, David Arbour, Aaron Mueller +2

Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating…

cs.LG2026

Training ML Models with Predictable Failures

Will Schwarzer, Scott Niekum

Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely large enough to observe the f…

cs.LG2026

Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control

Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee +1

Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost…

cs.LG2025

Adaptive Margin RLHF via Preference over Preferences

Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1

Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…

cs.LG2025

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Yaswanth Chittepu, Blossom Metevier, Will Schwarzer +3

Existing approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains. To ensure relia…