5 papers
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Haisen Luo, Yiwei Liu, Haoning Wang +13
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks. Existing self-distilla…
BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation
Zijun Jia, Yuanchang Ye, Sen Jia +6
Large language models (LLMs) can enhance factuality via retrieval-augmented generation (RAG), but applying RAG to every query is unnecessary when the model-only answer is reliable.…
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
Jiaming Zhang, Yujie Yang, Haoning Wang +2
Safe reinforcement learning (safe RL) aims to respect safety requirements while optimizing long-term performance. In many practical applications, however, the problem involves an i…
Low-rank Tensor Autoregressive Predictor for Third-Order Time-Series Forecasting
Haoning Wang, Liping Zhang
Recently, tensor time-series forecasting has gained increasing attention, whose core requirement is how to perform dimensionality reduction. In this paper, we establish a least squ…
Simplex Frank-Wolfe: Linear Convergence and Its Numerical Efficiency for Convex Optimization over Polytopes
Haoning Wang, Houduo Qi, Liping Zhang
We investigate variants of the Frank-Wolfe (FW) algorithm for smoothing and strongly convex optimization over polyhedral sets, with the goal of designing algorithms that achieve li…