Sim-To-Real Transfer for Visual Reinforcement Learning of Deformable Object Manipulation for Robot-Assisted Surgery
arXiv:2406.06092 · doi:10.1109/LRA.2022.3227873
Abstract
Automation holds the potential to assist surgeons in robotic interventions, shifting their mental work load from visuomotor control to high level decision making. Reinforcement learning has shown promising results in learning complex visuomotor policies, especially in simulation environments where many samples can be collected at low cost. A core challenge is learning policies in simulation that can be deployed in the real world, thereby overcoming the sim-to-real gap. In this work, we bridge the visual sim-to-real gap with an image-based reinforcement learning pipeline based on pixel-level domain adaptation and demonstrate its effectiveness on an image-based task in deformable object manipulation. We choose a tissue retraction task because of its importance in clinical reality of precise cancer surgery. After training in simulation on domain-translated images, our policy requires no retraining to perform tissue retraction with a 50% success rate on the real robotic system using raw RGB images. Furthermore, our sim-to-real transfer method makes no assumptions on the task itself and requires no paired images. This work introduces the first successful application of visual sim-to-real transfer for robotic manipulation of deformable objects in the surgical field, which represents a notable step towards the clinical translation of cognitive surgical robotics.
References in corpus (4)
- Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
- Curriculum Learning for Reinforcement Learning Domains: A Framework and Survey
- Autonomous Tissue Manipulation via Surgical Robot Using Learning Based Model Predictive Control
- Cooperative Assistance in Robotic Surgery through Multi-Agent Reinforcement Learning
Cited by in corpus (7)
- Movement Primitive Diffusion: Learning Gentle Robotic Manipulation of Deformable Objects
- From Decision to Action in Surgical Autonomy: Multi-Modal Large Language Models for Robot-Assisted Blood Suction
- FF-SRL: High Performance GPU-Based Surgical Simulation For Robot Learning
- Surgical Vision World Model
- Embedded Image-to-Image Translation for Efficient Sim-to-Real Transfer in Learning-based Robot-Assisted Soft Manipulation
- LUDO: Low-Latency Understanding of Deformable Objects using Point Cloud Occupancy Functions
- VERAGMIL: Virtual Environment for Scooping Granular Foods with Imitation Learning Models