5 papers
Test-Time Scaling via Error Localization
Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta +2
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standar…
REFA: Reference Free Alignment for multi-preference optimization
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2
To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effecti…
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +3
Direct Preference Optimization (DPO) has become a popular approach for aligning language models using pairwise preferences. However, in practical post-training pipelines, on-policy…
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2
Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, thereby enabling richer t…
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
Taneesh Gupta, Shivam Shandilya, Xuchao Zhang +5
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily…