activity
20242026
collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

GARDO: Reinforcing Diffusion Models without Reward Hacking

Haoran He, Yuxiao Ye, Jie Liu +7

Fine-tuning diffusion models via online reinforcement learning (RL) has shown great potential for enhancing text-to-image alignment. However, since precisely specifying a ground-tr…

cs.LG2025

Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards

Haoran He, Yuxiao Ye, Qingpeng Cai +4

RL with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving the reasoning abilities of large language models (LLMs). Current methods rely primarily on polic…

cs.LG2025

Random Policy Evaluation Uncovers Policies of Generative Flow Networks

Haoran He, Emmanuel Bengio, Qingpeng Cai +1

The Generative Flow Network (GFlowNet) is a probabilistic framework in which an agent learns a stochastic policy and flow functions to sample objects proportionally to an unnormali…

cs.LG2025

Looking Backward: Retrospective Backward Synthesis for Goal-Conditioned GFlowNets

Haoran He, Can Chang, Huazhe Xu +1

Generative Flow Networks (GFlowNets), a new family of probabilistic samplers, have demonstrated remarkable capabilities to generate diverse sets of high-reward candidates, in contr…

cs.LG2024

QGFN: Controllable Greediness with Action Values

Elaine Lau, Stephen Zhewen Lu, Ling Pan +2

Generative Flow Networks (GFlowNets; GFNs) are a family of energy-based generative methods for combinatorial objects, capable of generating diverse and high-utility samples. Howeve…

cs.LG2024

Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training

Haoran He, Chenjia Bai, Ling Pan +3

Learning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a…