34 citations · 63 across the 13 of their papers we have counts for
10 papers · 1 filter
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
Wenlong Deng, Yi Ren, Yushu Li +4
Reinforcement learning with verifiable rewards has significantly advanced the reasoning capabilities of large language models, yet how to explicitly steer training toward explorati…
Learning Dynamics of Deep Learning -- Force Analysis of Deep Neural Networks
Yi Ren
This thesis explores how deep learning models learn over time, using ideas inspired by force analysis. Specifically, we zoom in on the model's training procedure to see how one tra…
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
Wenlong Deng, Yi Ren, Muchen Li +3
Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a…
Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics
Yi Ren, Danica J. Sutherland
Obtaining compositional mappings is important for the model to generalize well compositionally. To better understand when and how to encourage the model to learn such mappings, we…
Learning Dynamics of LLM Finetuning
Yi Ren, Danica J. Sutherland
Learning dynamics, which describes how the learning of specific training examples influences the model's predictions on other examples, gives us a powerful tool for understanding t…
lpNTK: Better Generalisation with Less Data via Sample Interaction During Learning
Shangmin Guo, Yi Ren, Stefano V. Albrecht +1
Although much research has been done on proposing new models or loss functions to improve the generalisation of artificial neural networks (ANNs), less attention has been directed…