activity
20242026
collaborators

10 papers

cs.LG2026

It Takes Two: Your GRPO Is Secretly DPO

Yihong Wu, Liheng Ma, Lei Ding +9

GRPO has emerged as a prominent reinforcement learning algorithm for post-training LLMs. Unlike critic-based methods, GRPO computes advantages by estimating the \emph{value baselin…

cs.CL2026

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

Yihong Wu, Liheng Ma, Muzhi Li +7

Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to interface with search engines for co…

cs.LG2026

Plain Transformers Can be Powerful Graph Learners

Liheng Ma, Soumyasundar Pal, Yingxue Zhang +2

Transformers have attained outstanding performance across various modalities, owing to their simple but powerful scaled-dot-product (SDP) attention mechanisms. Researchers have att…

cs.LG2025

Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling

Derek Li, Jiaming Zhou, Leo Maxime Brunswic +8

The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-…

stat.ML2025

GraphPPD: Posterior Predictive Modelling for Graph-Level Inference

Soumyasundar Pal, Liheng Ma, Amine Natik +2

Accurate modelling and quantification of predictive uncertainty is crucial in deep learning since it allows a model to make safer decisions when the data is ambiguous and facilitat…

cs.LG2025

SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting

Yitian Zhang, Liheng Ma, Antonios Valkanas +2

Koopman operator theory provides a framework for nonlinear dynamical system analysis and time-series forecasting by mapping dynamics to a space of real-valued measurement functions…