3 papers
cs.LG2025
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
Haoran He, Yuxiao Ye, Qingpeng Cai +4
RL with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving the reasoning abilities of large language models (LLMs). Current methods rely primarily on polic…
cs.LG2025
Random Policy Evaluation Uncovers Policies of Generative Flow Networks
Haoran He, Emmanuel Bengio, Qingpeng Cai +1
The Generative Flow Network (GFlowNet) is a probabilistic framework in which an agent learns a stochastic policy and flow functions to sample objects proportionally to an unnormali…
cs.LG2024
Bifurcated Generative Flow Networks
Chunhui Li, Cheng-Hao Liu, Dianbo Liu +2
Generative Flow Networks (GFlowNets), a new family of probabilistic samplers, have recently emerged as a promising framework for learning stochastic policies that generate high-qua…