Publications (19)
Online Causal Kalman Filtering for Stable and Effective Policy Optimization
Shuo He, Lang Feng, Xin Cheng +2
Reinforcement learning for large language models suffers from high-variance token-level importance sampling (IS) ratios, which would destabilize policy optimization at scale. To im…
Test-Time Attention Purification for Backdoored Large Vision Language Models
Zhifang Zhang, Bojun Yang, Shuo He +5
Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…
Phase-Aware Mixture of Experts for Agentic Reinforcement Learning
Shengtian Yang, Yu Li, Shuo He +4
Reinforcement learning (RL) has equipped LLM agents with a strong ability to solve complex tasks. However, existing RL methods normally use a \emph{single} policy network, causing…
A Generalized Unbiased Risk Estimator for Learning with Augmented Classes
Senlin Shu, Shuo He, Haobo Wang +3
In contrast to the standard learning paradigm where all classes can be observed in training data, learning with augmented classes (LAC) tackles the problem where augmented classes…
Collaboration based Multi-Label Learning
Lei Feng, Bo An, Shuo He
It is well-known that exploiting label correlations is crucially important to multi-label learning. Most of the existing approaches take label correlations as prior knowledge, whic…
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
Lang Feng, Longtao Zheng, Shuo He +2
Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. I…
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
Shuo He, Lang Feng, Qi Wei +3
Group-based reinforcement learning (RL), such as GRPO, has advanced the capabilities of large language models on long-horizon agentic tasks. To enable more fine-grained policy upda…
Local Truncation Error-Guided Neural ODEs for Large Scale Traffic Forecasting
Xiao Zhang, Yafei Li, Ruixiang Wang +3
Spatiotemporal forecasting in physical systems, such as large-scale traffic networks, requires modeling a dual dynamic: continuous macroscopic rhythms and discrete, unpredictable m…
Test-Time Multimodal Backdoor Detection by Contrastive Prompting
Yuwei Niu, Shuo He, Qi Wei +3
While multimodal contrastive learning methods (e.g., CLIP) can achieve impressive zero-shot classification performance, recent research has revealed that these methods are vulnerab…
Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models
Zhifang Zhang, Jiahan Zhang, Shengjie Zhou +4
Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However,…
Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making
Guangfeng Cai, Kaibing Yang, Shuo He +4
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties super…
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
Kaibing Yang, Guangfeng Cai, Shengtian Yang +6
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.…
Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Xin Cheng, Shuo He, Lang Feng +4
Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agen…
Co-designed Quantum Discrete Adiabatic Linear System Solver Via Dynamic Circuits
Boxuan Ai, Shuo He, Xiang Zhao +7
Existing quantum discrete adiabatic approaches are hindered by circuit depth that increases linearly with the number of evolution steps, a significant challenge for current quantum…
Incorporating Multiple Cluster Centers for Multi-Label Learning
Senlin Shu, Fengmao Lv, Yan Yan +3
Multi-label learning deals with the problem that each instance is associated with multiple labels simultaneously. Most of the existing approaches aim to improve the performance of…
Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning
Zhifang Zhang, Shuo He, Haobo Wang +2
Multimodal contrastive learning models (e.g., CLIP) can learn high-quality representations from large-scale image-text datasets, while they exhibit significant vulnerabilities to b…
Partial-label Learning with Mixed Closed-set and Open-set Out-of-candidate Examples
Shuo He, Lei Feng, Guowu Yang
Partial-label learning (PLL) relies on a key assumption that the true label of each training example must be in the candidate label set. This restrictive assumption may be violated…
Compositional Complexity-Induced Ultralow Friction in Medium-Entropy MXenes
Jiaoli Li, Yuwei Zhang, Congjie Wei +11
Two-dimensional MXenes are promising solid lubricants, but the roles of compositional complexity and surface chemistry in governing interfacial friction remain unclear. Here, we sy…
Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation
Jianan Chen, Zhifang Zhang, Shuo He +3
Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of…