1 citations · 1 across the 14 of their papers we have counts for
5 papers · 1 filter
Ratio-Variance Regularized Policy Optimization
Yu Luo, Shuo Han, Yihan Hu +5
Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…
Theory-optimal Quantization Based on Flatness
Xiusheng Huang, Zhe Li, Xuanwu Yin +5
Post-training quantization has emerged as a widely adopted technique for compressing and accelerating the inference of Large Language Models (LLMs). The primary challenges in LLMs…
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
Yu Luo, Shuo Han, Yihan Hu +2
On-policy reinforcement learning (RL), particularly Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), has become the dominant paradigm for fine-tuni…
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
Derek Li, Jiaming Zhou, Leo Maxime Brunswic +8
The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-…
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time
Jikun Kang, Xin Zhe Li, Xi Chen +9
Although Large Language Models (LLMs) achieve remarkable performance across various tasks, they often struggle with complex reasoning tasks, such as answering mathematical question…