From the 1 of 7 linked papers with an AI index.
1 citations · 1 across the 7 of their papers we have counts for
5 papers · 1 filter
Distilled Reinforcement Learning for LLM Post-training
Chen Wang, Zhaochun Li, Jionghao Bai +4
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL)…
Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment
Chunyu Hu, Tianyin Liao, Ge Lan +4
The paper introduces GTAlign, a simple framework that aligns graph structures to tabular representations, enabling a text-free Graph Foundation Model that uses community-guided con…
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Chen Wang, Zhaochun Li, Jionghao Bai +3
Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entro…
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
Zhaochun Li, Chen Wang, Jionghao Bai +4
The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…
Unlocking the Potentials of Retrieval-Augmented Generation for Diffusion Language Models
Chuanyue Yu, Jiahui Wang, Yuhan Li +6
Diffusion Language Models (DLMs) have recently demonstrated remarkable capabilities in natural language processing tasks. However, the potential of Retrieval-Augmented Generation (…