collaborators

8 papers

cs.LG2026

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Tao Wang, Shuo Li, Yan Sun +2

Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimi…

cs.LG2026

The Forecast After the Forecast: A Post-Processing Shift in Time Series

Daojun Liang, Qi Li, Yinglong Wang +5

Time series forecasting has long been dominated by advances in model architecture, with recent progress driven by deep learning and hybrid statistical techniques. However, as forec…

cs.AI2025

Shylock: Causal Discovery in Multivariate Time Series based on Hybrid Constraints

Shuo Li, Keqin Xu, Jie Liu +1

Causal relationship discovery has been drawing increasing attention due to its prevalent application. Existing methods rely on human experience, statistical methods, or graphical c…

cs.LG2025

RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models

Lianghuan Huang, Sagnik Anupam, Insup Lee +2

Reinforcement learning (RL) has emerged as a promising strategy for finetuning small language models (SLMs) to solve targeted tasks such as math and coding. However, RL algorithms…

cs.CL2025

MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety

Yahan Yang, Soham Dan, Shuo Li +2

Large Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors. This vulnerability is exacerbated in multili…

cs.LG2025

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

Haosen Ge, Shuo Li, Lianghuan Huang

Effective prompt engineering remains a challenging task for many applications. We introduce Weak-to-Strong Transfer (WST), an automatic prompt engineering framework where a small "…