papers

Publications (8)

cs.CL2025

CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling

Taneesh Gupta, Shivam Shandilya, Xuchao Zhang +5

Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily…

cs.LG2026

Test-Time Scaling via Error Localization

Rajiv Shailesh Chitale, Rahul Madhavan, Taneesh Gupta +2

Scaling inference-time computation has emerged as a reliable method to improve the performance of large language models on complex reasoning and programming tasks. However, standar…

cs.LG2025

REFA: Reference Free Alignment for multi-preference optimization

Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2

To mitigate reward hacking from response verbosity, modern preference optimization methods are increasingly adopting length normalization (e.g., SimPO, ORPO, LN-DPO). While effecti…

cs.CV2024

Map It Anywhere (MIA): Empowering Bird's Eye View Mapping using Large-scale Public Data

Cherie Ho, Jiaye Zou, Omar Alama +7

Top-down Bird's Eye View (BEV) maps are a popular representation for ground robot navigation due to their richness and flexibility for downstream tasks. While recent methods have s…

cs.LG2025

Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts

Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +3

Direct Preference Optimization (DPO) has become a popular approach for aligning language models using pairwise preferences. However, in practical post-training pipelines, on-policy…

cs.LG2025

AMPO: Active Multi-Preference Optimization for Self-play Preference Selection

Taneesh Gupta, Rahul Madhavan, Xuchao Zhang +2

Multi-preference optimization enriches language-model alignment beyond pairwise preferences by contrasting entire sets of helpful and undesired responses, thereby enabling richer t…