8 papers
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
Tao Wang, Shuo Li, Yan Sun +2
Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimi…
The Forecast After the Forecast: A Post-Processing Shift in Time Series
Daojun Liang, Qi Li, Yinglong Wang +5
Time series forecasting has long been dominated by advances in model architecture, with recent progress driven by deep learning and hybrid statistical techniques. However, as forec…
Shylock: Causal Discovery in Multivariate Time Series based on Hybrid Constraints
Shuo Li, Keqin Xu, Jie Liu +1
Causal relationship discovery has been drawing increasing attention due to its prevalent application. Existing methods rely on human experience, statistical methods, or graphical c…
RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
Lianghuan Huang, Sagnik Anupam, Insup Lee +2
Reinforcement learning (RL) has emerged as a promising strategy for finetuning small language models (SLMs) to solve targeted tasks such as math and coding. However, RL algorithms…
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
Yahan Yang, Soham Dan, Shuo Li +2
Large Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors. This vulnerability is exacerbated in multili…
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
Haosen Ge, Shuo Li, Lianghuan Huang
Effective prompt engineering remains a challenging task for many applications. We introduce Weak-to-Strong Transfer (WST), an automatic prompt engineering framework where a small "…