3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.LG2025
RLSR: Reinforcement Learning from Self Reward
Toby Simonds, Kevin Lopez, Akira Yoshiyama +1
Large language models can generate solutions to complex problems, but training them with reinforcement learning typically requires verifiable rewards that are expensive to create a…
cs.LG2025★ 3 cited
LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
Toby Simonds, Akira Yoshiyama
We introduce LADDER (Learning through Autonomous Difficulty-Driven Example Recursion), a framework which enables Large Language Models to autonomously improve their problem-solving…