2 papers
cs.CV2026
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
Liyu Zhang, Kehan Li, Tingrui Han +4
Post training via GRPO has demonstrated remarkable effectiveness in improving the generation quality of flow-matching models. However, GRPO suffers from inherently low sample effic…
cs.CL2026
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
Ran Xu, Tianci Liu, Zihan Dong +6
Standard reward models typically predict scalar scores that fail to capture the multifaceted nature of response quality in non-verifiable domains, such as creative writing or open-…