3 citations · 8 across the 18 of their papers we have counts for
22 papers
Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning
Zhenya Liu, Yang Meng, Zhuokai Zhao +2
When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations…
Rethinking Transfer in Continual Learning: A Replay-Based Realisation
Yang Meng, Zhenya Liu, Zhuokai Zhao +1
Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (repl…
Accelerating PDE Surrogates via RL-Guided Mesh Optimization
Yang Meng, Ruoxi Jiang, Zhuokai Zhao +3
Deep surrogate models for parametric partial differential equations (PDEs) can deliver high-fidelity approximations but remain prohibitively data-hungry: training often requires th…
Scaling Agent Learning via Experience Synthesis
Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15
While reinforcement learning (RL) can empower autonomous agents by enabling self-improvement through interaction, its practical adoption remains challenging due to costly rollouts,…
Boosting LLM Reasoning via Spontaneous Self-Correction
Xutong Zhao, Tengyu Xu, Xuewei Wang +11
While large language models (LLMs) have demonstrated remarkable success on a broad range of tasks, math reasoning remains a challenging one. One of the approaches for improving mat…
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
Yuhang Zhou, Jing Zhu, Shengyi Qian +7
Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Op…