From the 1 of 7 linked papers with an AI index.
7 papers
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Manas Pathak, Xingyao Chen, Shuozhe Li +2
The paper introduces the Filtered Reasoning Score (FRS), a metric that evaluates the quality of reasoning traces from large language models by focusing on the most confident genera…
-Balancing for Mixture-of-Experts Training
Lizhang Chen, Jonathan Li, Qi Wang +5
Mixture-of-Experts (MoE) models rely on balanced expert utilization to fully realize their scalability. However, existing load-balancing methods are largely heuristic and operate o…
A Learnable Wavelet Transformer for Long-Short Equity Trading and Risk-Adjusted Return Optimization
Shuozhe Li, Du Cheng, Leqi Liu
Learning profitable intraday trading policies from financial time series is challenging due to heavy noise, non-stationarity, and strong cross-sectional dependence among related as…
Learning Robust Reasoning through Guided Adversarial Self-Play
Shuozhe Li, Vaishnav Tadiparthi, Kwonjoon Lee +6
Reinforcement learning from verifiable rewards (RLVR) produces strong reasoning models, yet they can fail catastrophically when the conditioning context is fallible (e.g., corrupte…
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
Ruiyang Zhou, Shuozhe Li, Amy Zhang +1
Self-improvement via RL often fails on complex reasoning tasks because GRPO-style post-training methods rely on the model's initial ability to generate positive samples. Without gu…
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
Shuozhe Li, Jincheng Cao, Bodun Hu +3
Reinforcement finetuning (RFT) has emerged as a powerful paradigm for unlocking reasoning capabilities in large language models. However, we identify a critical trade-off: while un…