collaborators

5 papers

stat.ME2026

Denoising Data with Measurement Error Using a Reproducing Kernel-based Diffusion Model

Mingyang Yi, Marcos Matabuena, Zhiming Ma +1

The ongoing technological revolution in measurement systems enables the acquisition of high-resolution samples in fields such as engineering, biology, and medicine. However, these…

cs.LG2025

Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Xuerui Su, Shufang Xie, Guoqing Liu +7

Recently, Large Language Models (LLMs) have rapidly evolved, approaching Artificial General Intelligence (AGI) while benefiting from large-scale reinforcement learning to enhance H…

cs.LG2025

DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

Xuerui Su, Liya Guo, Yue Wang +4

Inference scaling further accelerates Large Language Models (LLMs) toward Artificial General Intelligence (AGI), with large-scale Reinforcement Learning (RL) to unleash long Chain-…

eess.SY2025

Model-Based Closed-Loop Control Algorithm for Stochastic Partial Differential Equation Control

Peiyan Hu, Haodong Feng, Yue Wang +1

Neural operators have demonstrated promise in modeling and controlling systems governed by Partial Differential Equations (PDEs). Beyond PDEs, Stochastic Partial Differential Equat…

cs.LG2025

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms

Xuerui Su, Yue Wang, Jinhua Zhu +4

With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and a…