activity
20242026
collaborators

8 papers

stat.ML2026

Optimal Mixture-of-Experts Model Averaging for Conditional Generative Models

Shijin Gong, Baihua He, Xinyu Zhang

Conditional generative models have emerged as powerful tools for sampling from target conditional distributions, driving substantial advances across a wide range of scientific and…

cs.LG2026

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

Shijin Gong, Erhan Xu, Kai Ye +3

Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff betw…

cs.CL2026

READER: Reasoning-Enhanced AI-Generated Text Detection

Pingfan Su, Kai Ye, Shijin Gong +4

Recent advances in large language models (LLMs) have made it increasingly difficult to distinguish human-written text from AI-generated content. Many existing detectors train super…

cs.LG2026

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning

Shijin Gong, Kai Ye, Jin Zhu +3

Recent advances in large language models (LLMs) have increasingly relied on reinforcement learning (RL) to improve their reasoning capabilities. Three types of approaches have been…

stat.ME2026

Synthetic Control Method with Mixed Frequency Data

Lu Zhang, Shijin Gong, Xinyu Zhang

Mixed-frequency data, where variables are observed at different temporal resolutions, commonly occur in economic and financial studies. Classical synthetic control methods (SCM) ar…

cs.LG2026

Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic

Hongyi Zhou, Kai Ye, Erhan Xu +4

Group relative policy optimization (GRPO), a core methodological component of DeepSeekMath and DeepSeek-R1, has emerged as a cornerstone for scaling reasoning capabilities of large…