activity
20192026
most citedA Hierarchical Transformer with Speaker Modeling for Emotion Recognition in Conversation

11 citations · 23 across the 44 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

Naibin Gu, Qingyi Si, Chenxu Yang +5

On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same mo…

cs.LG2026

Learning to Solve, Forgetting to Retain: Correct-Set Turnover in RLVR

Chuanyu Qin, Chenxu Yang, Qingyi Si +3

Reinforcement learning with verifiable rewards (RLVR) improves the ability of large language model, yet headline accuracy gains often conceal a hidden cost: previously solved probl…

cs.LG2026

Co-Evolving Policy Distillation

Naibin Gu, Chenxu Yang, Qingyi Si +7

RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single mode…

cs.LG2026

Near-Future Policy Optimization

Chuanyu Qin, Chenxu Yang, Qingyi Si +6

Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…

cs.LG2026

Self-Distilled RLVR

Chenxu Yang, Chuanyu Qin, Qingyi Si +7

On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals…

cs.LG2026

IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models

Haonan Song, Qingchen Xie, Huan Zhu +12

Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (…