3 papers
cs.LG2025
Robust Reward Modeling via Causal Rubrics
Pragya Srivastava, Harman Singh, Rahul Madhavan +9
Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or…
cs.CL2025
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
Sravanti Addepalli, Yerram Varun, Arun Suggala +2
Large Language Models (LLMs) are known to be susceptible to crafted adversarial attacks or jailbreaks that lead to the generation of objectionable content despite being aligned to…
cs.CL2025
Time-Reversal Provides Unsupervised Feedback to LLMs
Yerram Varun, Rahul Madhavan, Sravanti Addepalli +3
Large Language Models (LLMs) are typically trained to predict in the forward direction of time. However, recent works have shown that prompting these models to look back and critiq…