5 papers
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
Branislav Kveton, Anup Rao, Subhojyoti Mukherjee +2
We introduce AdvantageFlow, a forward-process reinforcement learning algorithm for rectified flow models. Unlike Flow-GRPO, which optimizes the reverse process, we optimize an adva…
Partial Policy Gradients for RL in LLMs
Puneet Mathur, Branislav Kveton, Subhojyoti Mukherjee +1
Reinforcement learning is a framework for learning to act sequentially in an unknown environment. We propose a natural approach for modeling policy structure in policy gradients. T…
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
Thanh-Long V. Le, Myeongho Jeon, Kim Vu +2
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful framework for improving the reasoning abilities of Large Language Models (LLMs). However, current methods such a…
DocFinQA: A Long-Context Financial Reasoning Dataset
Varshini Reddy, Rik Koncel-Kedziorski, Viet Dac Lai +3
For large language models (LLMs) to be effective in the financial domain -- where each decision can have a significant impact -- it is necessary to investigate realistic tasks and…
Language Model Probabilities are Not Calibrated in Numeric Contexts
Charles Lovering, Michael Krumdick, Viet Dac Lai +5
Some statements have one well-defined continuation (e.g., "the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., "the weighted c…