activity
20232026
most citedHALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

3 citations · 8 across the 18 of their papers we have counts for

collaborators

22 papers

cs.LG2026

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

Zhenya Liu, Yang Meng, Zhuokai Zhao +2

When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations…

cs.LG2026

Rethinking Transfer in Continual Learning: A Replay-Based Realisation

Yang Meng, Zhenya Liu, Zhuokai Zhao +1

Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch. Existing methods, whether rehearsal-based (repl…

cs.LG2026

Accelerating PDE Surrogates via RL-Guided Mesh Optimization

Yang Meng, Ruoxi Jiang, Zhuokai Zhao +3

Deep surrogate models for parametric partial differential equations (PDEs) can deliver high-fidelity approximations but remain prohibitively data-hungry: training often requires th…

cs.AI2025

Scaling Agent Learning via Experience Synthesis

Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15

While reinforcement learning (RL) can empower autonomous agents by enabling self-improvement through interaction, its practical adoption remains challenging due to costly rollouts,…

cs.AI2025

Boosting LLM Reasoning via Spontaneous Self-Correction

Xutong Zhao, Tengyu Xu, Xuewei Wang +11

While large language models (LLMs) have demonstrated remarkable success on a broad range of tasks, math reasoning remains a challenging one. One of the approaches for improving mat…

cs.CL2025

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

Yuhang Zhou, Jing Zhu, Shengyi Qian +7

Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Op…