2 papers
cs.LG2026
Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning
Alan Nadelsticher Ruvalcaba
The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational p…
cs.LG2026
When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction
Tyler Crosse, Alan Nadelsticher Ruvalcaba, Dustin Khang LeDuc +3
Different predictors often excel on different inputs, so picking the best one per instance promises higher accuracy than committing to a single model. In practice, selectors traine…