4 papers
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Ankur Samanta, Akshayaa Magesh, Tal Lancewicki +7
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic uncertainty about their environm…
Formalizing Learning from Language Feedback with Provable Guarantees
Wanqiao Xu, Allen Nie, Ruijie Zheng +3
Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents. Despite impressive empir…
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
Ying Fan, Jingling Li, Adith Swaminathan +2
We present a novel method, Contextual goal-Oriented Data Augmentation (CODA), which uses commonly available unlabeled trajectories and context-goal pairs to solve Contextual Goal-O…