Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +4
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…
cs.LG2026
Formalizing Learning from Language Feedback with Provable Guarantees
Wanqiao Xu, Allen Nie, Ruijie Zheng +3
Interactively learning from observation and language feedback is an increasingly studied area driven by the emergence of large language model (LLM) agents. Despite impressive empir…
cs.LG2025
How to Solve Contextual Goal-Oriented Problems with Offline Datasets?
Ying Fan, Jingling Li, Adith Swaminathan +2
We present a novel method, Contextual goal-Oriented Data Augmentation (CODA), which uses commonly available unlabeled trajectories and context-goal pairs to solve Contextual Goal-O…