2 papers
cs.LG2026
Gradient Extrapolation-Based Policy Optimization
Ismam Nur Swapnil, Aranya Saha, Tanvir Ahmed Khan +2
Reinforcement learning is widely used to improve the reasoning ability of large language models, especially when answers can be automatically checked. Standard GRPO-style training…
cs.LG2024
Superpipeline: A Universal Approach for Reducing GPU Memory Usage in Large Models
Reza Abbasi, Sernam Lim
The rapid growth in machine learning models, especially in natural language processing and computer vision, has led to challenges when running these models on hardware with limited…