papers

Publications (19)

cs.CL2026

Online Causal Kalman Filtering for Stable and Effective Policy Optimization

Shuo He, Lang Feng, Xin Cheng +2

Reinforcement learning for large language models suffers from high-variance token-level importance sampling (IS) ratios, which would destabilize policy optimization at scale. To im…

cs.CV2026

Test-Time Attention Purification for Backdoored Large Vision Language Models

Zhifang Zhang, Bojun Yang, Shuo He +5

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…

cs.AI2026

Phase-Aware Mixture of Experts for Agentic Reinforcement Learning

Shengtian Yang, Yu Li, Shuo He +4

Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…

cs.LG2023

A Generalized Unbiased Risk Estimator for Learning with Augmented Classes

Senlin Shu, Shuo He, Haobo Wang +3

In contrast to the standard learning paradigm where all classes can be observed in training data, learning with augmented classes (LAC) tackles the problem where augmented classes…

cs.LG2019

Collaboration based Multi-Label Learning

Lei Feng, Bo An, Shuo He

It is well-known that exploiting label correlations is crucially important to multi-label learning. Most of the existing approaches take label correlations as prior knowledge, whic…

cs.LG2026

Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems

Lang Feng, Longtao Zheng, Shuo He +2

Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. I…

cs.LG2026

Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks

Shuo He, Lang Feng, Qi Wei +3

Group-based reinforcement learning (RL), such as GRPO, has advanced the capabilities of large language models on long-horizon agentic tasks. To enable more fine-grained policy upda…

cs.LG2026

Local Truncation Error-Guided Neural ODEs for Large Scale Traffic Forecasting

Xiao Zhang, Yafei Li, Ruixiang Wang +3

Spatiotemporal forecasting in physical systems, such as large-scale traffic networks, requires modeling a dual dynamic: continuous macroscopic rhythms and discrete, unpredictable m…

cs.CV2025

Test-Time Multimodal Backdoor Detection by Contrastive Prompting

Yuwei Niu, Shuo He, Qi Wei +3

While multimodal contrastive learning methods (e.g., CLIP) can achieve impressive zero-shot classification performance, recent research has revealed that these methods are vulnerab…

cs.CV2025

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

Zhifang Zhang, Jiahan Zhang, Shengjie Zhou +4

Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However,…

cs.CL2026

Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making

Guangfeng Cai, Kaibing Yang, Shuo He +4

Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…

cs.LG2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Kaibing Yang, Guangfeng Cai, Shengtian Yang +6

Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…

cs.LG2026

Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

Xin Cheng, Shuo He, Lang Feng +4

Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agen…

quant-ph2025

Co-designed Quantum Discrete Adiabatic Linear System Solver Via Dynamic Circuits

Boxuan Ai, Shuo He, Xiang Zhao +7

Existing quantum discrete adiabatic approaches are hindered by circuit depth that increases linearly with the number of evolution steps, a significant challenge for current quantum…

cs.LG2022

Incorporating Multiple Cluster Centers for Multi-Label Learning

Senlin Shu, Fengmao Lv, Yan Yan +3

Multi-label learning deals with the problem that each instance is associated with multiple labels simultaneously. Most of the existing approaches aim to improve the performance of…

cs.CV2025

Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning

Zhifang Zhang, Shuo He, Haobo Wang +2

Multimodal contrastive learning models (e.g., CLIP) can learn high-quality representations from large-scale image-text datasets, while they exhibit significant vulnerabilities to b…

cs.LG2023

Partial-label Learning with Mixed Closed-set and Open-set Out-of-candidate Examples

Shuo He, Lei Feng, Guowu Yang

Partial-label learning (PLL) relies on a key assumption that the true label of each training example must be in the candidate label set. This restrictive assumption may be violated…

cond-mat.mtrl-sci2026

Compositional Complexity-Induced Ultralow Friction in Medium-Entropy MXenes

Jiaoli Li, Yuwei Zhang, Congjie Wei +11

Two-dimensional MXenes are promising solid lubricants, but the roles of compositional complexity and surface chemistry in governing interfacial friction remain unclear. Here, we sy…

cs.AI2026

Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation

Jianan Chen, Zhifang Zhang, Shuo He +3

Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of…