Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
RLSR: Reinforcement Learning from Self Reward
Toby Simonds, Kevin Lopez, Akira Yoshiyama +1
Large language models can generate solutions to complex problems, but training them with reinforcement learning typically requires verifiable rewards that are expensive to create a…
cs.LG2025
LADDER: Self-Improving LLMs Through Recursive Problem Decomposition
Toby Simonds, Akira Yoshiyama
We introduce LADDER (Learning through Autonomous Difficulty-Driven Example Recursion), a framework which enables Large Language Models to autonomously improve their problem-solving…
cs.LG2025
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
Toby Simonds
We present Entropy Adaptive Decoding (EAD), a novel approach for efficient language model inference that dynamically switches between different-sized models based on prediction unc…