3 papers
cs.LG2026
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL
Huy Le, Onur Celik, Denis Blessing +6
Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely…
cs.LG2026
PAWS: Preference Learning with Advantage-Weighted Segments
Aleksandar Taranovic, Onur Celik, Niklas Freymuth +6
Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods…
cs.RO2022
Autonomous soft hand grasping -- Literature review
Tai Hoang
Autonomous grasping remains challenging as unlike humans, robots do not possess a sophisticated sensing nor delicate interaction capability with the real environment. Among other e…