3 papers
cs.LG2026
A Geometric Perspective on Stabilizing Value Conflict Resolution
Saket Reddy, Andy Liu
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To add…
cs.LG2026
ShallowBench: Benchmarking Generative Drug Design Models on Shallow-Pocket Targets
Saket Reddy, Shiwei Liu
While generative AI models have demonstrated remarkable success in structure-based drug design, they predominantly rely on deep binding pockets and struggle to sample effective lig…
cs.AI2026
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
Saket Reddy, Ke Yang, ChengXiang Zhai
Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground truth, creating a high-variance,…