Publications (8)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
Taneesh Gupta, Shivam Shandilya, Xuchao Zhang +5
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily…
Test-Time Scaling via Error Localization
Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta +2
Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standar…
REFA: Reference Free Alignment for multi-preference optimization
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2
To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effecti…
Map It Anywhere (MIA): Empowering Bird's Eye View Mapping using Large-scale Public Data
Cherie Ho, Jiaye Zou, Omar Alama +7
Top-down Bird's Eye View (BEV) maps are a popular representation for ground robot navigation due to their richness and flexibility for downstream tasks. While recent methods have s…
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +3
Direct Preference Optimization (DPO) has become a popular approach for aligning language models using pairwise preferences. However, in practical post-training pipelines, on-policy…
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2
Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, thereby enabling richer t…