3 papers
cs.NE2026
Reinforcement Learning from Meta-Evaluation: Aligning Language Models Without Ground-Truth Labels
Micah Rentschler, Jesse Roberts
Most reinforcement learning (RL) methods for training large language models (LLMs) require ground-truth labels or task-specific verifiers, limiting scalability when correctness is…
cs.LG2025
Exploitation Is All You Need... for Exploration
Micah Rentschler, Jesse Roberts
Ensuring sufficient exploration is a central challenge when training meta-reinforcement learning (meta-RL) agents to solve novel environments. Conventional solutions to the explora…
cs.LG2025
RL + Transformer = A General-Purpose Problem Solver
Micah Rentschler, Jesse Roberts
What if artificial intelligence could not only solve problems for which it was trained but also learn to teach itself to solve new problems (i.e., meta-learn)? In this study, we de…