5 papers
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
Soham Sane
Proximal Policy Optimization (PPO) is a widely used reinforcement learning algorithm that heavily relies on accurate advantage estimates for stable and efficient training. However,…
AlphaGrad: Non-Linear Gradient Normalization Optimizer
Soham Sane
We introduce AlphaGrad, a memory-efficient, conditionally stateless optimizer addressing the memory overhead and hyperparameter complexity of adaptive methods like Adam. AlphaGrad…
Optimal Scaling Laws for Efficiency Gains in a Theoretical Transformer-Augmented Sectional MoE Framework
Soham Sane
This paper introduces a theoretical framework for a Transformer-augmented, sectional Mixture-of-Experts (MoE) architecture that aims to enhance computational efficiency while prese…
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
Soham Sane
Hybrid Group Relative Policy Optimization (Hybrid GRPO) is a reinforcement learning framework that extends Proximal Policy Optimization (PPO) and Group Relative Policy Optimization…
A NotSo Simple Way to Beat Simple Bench
Soham Sane, Angus McLean
This paper presents a novel framework for enhancing reasoning capabilities in large language models (LLMs) by leveraging iterative reasoning and feedback-driven methodologies. Buil…