2 papers
cs.AI2025
Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
Arpan Mukherjee, Marcello Bullo, Debabrota Basu +1
While test-time scaling with verification has shown promise in improving the performance of large language models (LLMs), the role of the verifier and its imperfections remain unde…
cs.LG2025
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
Arpan Mukherjee, Marcello Bullo, Deniz Gündüz
Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity…