3 papers
cs.AI2026
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
Anisha Garg, Claire Zhang, Nishit Neema +3
Group-Relative Policy Optimization (GRPO) has emerged as the standard for training reasoning capabilities in large language models through reinforcement learning. By estimating adv…
cs.CL2025
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
Nishit Neema, Srinjoy Mukherjee, Sapan Shah +2
Large Language Models (LLMs) excel at general tasks but underperform in specialized domains like economics and psychology, which require deep, principled understanding. To address…
cs.AI2025
Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
Anisha Garg, Engin Tekin, Yash More +3
Advanced test-time computing strategies are essential for scaling reasoning models, but their effectiveness is capped by the models' poor self-evaluation. We propose a pairwise Exp…