Publications (12)
Gaze patterns predict preference and confidence in pairwise AI image evaluation
Nikolas Papadopoulos, Shreenithi Navaneethan, Sheng Bai +2
Preference learning methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on pairwise human judgments, yet little is kno…
Structure Enables Effective Self-Localization of Errors in LLMs
Ankur Samanta, Akshayaa Magesh, Ayush Jain +8
Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect reasoning, as a path toward buildin…
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
Runzhe Wu, Ankur Samanta, Ayush Jain +7
Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environm…
FragmentNet: Adaptive Graph Fragmentation for Graph-to-Sequence Molecular Representation Learning
Ankur Samanta, Rohan Gupta, Aditi Misra +2
Molecular representation learning methods typically tokenize molecules as individual atoms or use rigid, rule-based fragment decompositions, limiting their ability to capture meani…
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
Amit Roth, Ankur Samanta, Matan Halevy +2
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…