most citedRLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

10 citations · 10 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback

Shreyas Chaudhari, Renhao Zhang, Philip S. Thomas +1

The ability of reinforcement learning algorithms to learn effective policies is determined by the rewards available during training. However, for practical problems, obtaining larg…

cs.CL2025

Probing AI Safety with Source Code

Ujwal Narayan, Shreyas Chaudhari, Ashwin Kalyan +4

Large language models (LLMs) have become ubiquitous, interfacing with humans in numerous safety-critical applications. This necessitates improving capabilities, but importantly cou…

cs.AI2025

Agent Context Protocols Enhance Collective Inference

Devansh Bhardwaj, Arjun Beniwal, Shreyas Chaudhari +5

AI agents have become increasingly adept at complex tasks such as coding, reasoning, and multimodal understanding. However, building generalist systems requires moving beyond indiv…

cs.LG2025

ReLU Networks as Random Functions: Their Distribution in Probability Space

Shreyas Chaudhari, José M. F. Moura

This paper presents a novel framework for understanding trained ReLU networks as random, affine functions, where the randomness is induced by the distribution over the inputs. By c…

cs.LG202410 cited

RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari +5

State-of-the-art large language models (LLMs) have become indispensable tools for various tasks. However, training LLMs to serve as effective assistants for humans requires careful…