2 papers
cs.LG2026
Soft : A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
Pranav Mahajan, Ben Seymour
Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a penalty on the divergence from…
cs.RO2025
Neural Associative Skill Memories for safer robotics and modelling human sensorimotor repertoires
Pranav Mahajan, Mufeng Tang, T. Ed Li +2
Modern robots face challenges shared by humans, where machines must learn multiple sensorimotor skills and express them adaptively. Equipping robots with a human-like memory of how…