8 papers
Optimal Mixture-of-Experts Model Averaging for Conditional Generative Models
Shijin Gong, Baihua He, Xinyu Zhang
Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and…
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
Shijin Gong, Erhan Xu, Kai Ye +3
Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff betw…
READER: Reasoning-Enhanced AI-Generated Text Detection
Pingfan Su, Kai Ye, Shijin Gong +4
Recent advances in large language models (LLMs) have made it increasingly difficult to distinguish human-written text from AI-generated content. Many existing detectors train super…
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning
Shijin Gong, Kai Ye, Jin Zhu +3
Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three types of approaches have been…
Synthetic Control Method with Mixed Frequency Data
Lu Zhang, Shijin Gong, Xinyu Zhang
Mixed-frequency data, where variables are observed at different temporal resolutions, commonly occur in economic and financial studies. Classical synthetic control methods (SCM) ar…
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
Hongyi Zhou, Kai Ye, Erhan Xu +4
Group relative policy optimization (GRPO), a core methodological component of DeepSeekMath and DeepSeek-R1, has emerged as a cornerstone for scaling reasoning capabilities of large…