activity
20242026
collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

Kai Yang, Jingwei Xu, Wanyu Wang +4

On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradat…

cs.LG2026

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex

Yun Qu, Qi Wang, Yixiu Mao +11

Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning capacity. Among existin…

cs.LG2026

Debiased Model-based Representations for Sample-efficient Continuous Control

Jiafei Lyu, Zichuan Lin, Scott Fujimoto +5

Model-based representations recently stand out as a promising framework that embeds latent dynamics information into the representations for downstream off-policy actor-critic lear…

cs.LG2025

Exploration by Random Distribution Distillation

Zhirui Fang, Kai Yang, Jian Tao +4

Exploration remains a critical challenge in online reinforcement learning, as an agent must effectively explore unknown environments to achieve high returns. Currently, the main ex…

cs.LG2024

Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning

Yangkun Chen, Kai Yang, Jian Tao +1

Recently, deep Multi-Agent Reinforcement Learning (MARL) has demonstrated its potential to tackle complex cooperative tasks, pushing the boundaries of AI in collaborative environme…

cs.LG2024

A Two-stage Reinforcement Learning-based Approach for Multi-entity Task Allocation

Aicheng Gong, Kai Yang, Jiafei Lyu +1

Task allocation is a key combinatorial optimization problem, crucial for modern applications such as multi-robot cooperation and resource scheduling. Decision makers must allocate…