3 papers
stat.ML2026
Provably Reliable Classifier Guidance via Cross-Entropy Control
Sharan Sahu, Arisina Banerjee, Yuchen Wu
Classifier-guided diffusion models generate conditional samples by augmenting the reverse-time score with the gradient of the log-probability predicted by a probabilistic classifie…
cs.LG2025
Online Distributionally Robust LLM Alignment via Regression to Relative Reward
Sharan Sahu, Martin T. Wells
Reinforcement Learning with Human Feedback (RLHF) has become crucial for aligning Large Language Models (LLMs) with human intent. However, existing offline RLHF approaches suffer f…
cs.LG2025
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
Sharan Sahu
We study regret minimization under privacy constraints in episodic inhomogeneous linear Markov Decision Processes (MDPs), motivated by the growing use of reinforcement learning (RL…