3 papers
cs.RO2026
FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation
Kexin Shi, Junyao Shi, Poorvi Hebbar +5
Real-world reinforcement learning for robotic manipulation remains challenging, and this difficulty is amplified for flow matching policies: applying policy gradient methods to the…
cs.LG2025
Evolutionary Policy Optimization
Jianren Wang, Yifan Su, Abhinav Gupta +1
On-policy reinforcement learning (RL) algorithms are widely used for their strong asymptotic performance and training stability, but they struggle to scale with larger batch sizes,…
cs.CL2025
ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
Hao Wen, Yifan Su, Feifei Zhang +4
Recent advances in Large Language Models (LLMs) have been driven by test-time compute scaling - a strategy that improves reasoning by generating longer, sequential thought processe…