3 papers
cs.LG2026
Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
Qi Wang, Mian Wu, Yuyang Zhang +7
Reinforcement Learning (RL) has achieved remarkable success in various domains, yet it often relies on carefully designed programmatic reward functions to guide agent behavior. Des…
cs.DM2026
A Scalable Lift-and-Project Differentiable Approach For the Maximum Cut Problem
Ismail Alkhouri, Mian Wu, Cunxi Yu +3
We propose a scalable framework for solving the Maximum Cut (MaxCut) problem in large graphs using projected gradient ascent on quadratic objectives. Our approach is differentiable…
cs.LG2025
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
Mian Wu, Gavin Zhang, Sewon Min +2
Open-ended generation tasks require outputs to satisfy diverse and often implicit task-specific evaluation rubrics. The sheer number of relevant rubrics leads to prohibitively high…