4 papers
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
Subramanyam Sahoo, Vinija Jain, Saanidhya Vats +4
Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a…
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
Subramanyam Sahoo, Aman Chadha, Vinija Jain +1
Reinforcement Learning from Human Feedback (RLHF) is widely used for aligning large language models, yet practitioners face a persistent puzzle: improving safety often reduces fair…
The Last Vote: A Multi-Stakeholder Framework for Language Model Governance
Subramanyam Sahoo, Aditi Chhawacharia
As artificial intelligence systems become increasingly powerful and pervasive, democratic societies face unprecedented challenges in governing these technologies while preserving c…
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
Subramanyam Sahoo
Reward design is central to reinforcement learning from human feedback (RLHF) and alignment research. In this work, we propose a unified framework to study hard, continuous, and hy…