A Survey of Embodied Learning for Object-Centric Robotic Manipulation
arXiv:2408.11537 · doi:10.1007/s11633-025-1542-8
Abstract
Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike data-driven machine learning methods, embodied learning focuses on robot learning through physical interaction with the environment and perceptual feedback, making it especially suitable for robotic manipulation. In this paper, we provide a comprehensive survey of the latest advancements in this field and categorize the existing work into three main branches: 1) Embodied perceptual learning, which aims to predict object pose and affordance through various data representations; 2) Embodied policy learning, which focuses on generating optimal robotic decisions using methods such as reinforcement learning and imitation learning; 3) Embodied task-oriented learning, designed to optimize the robot's performance based on the characteristics of different tasks in object grasping and manipulation. In addition, we offer an overview and discussion of public datasets, evaluation metrics, representative applications, current challenges, and potential future research directions. A project associated with this survey has been established at https://github.com/RayYoh/OCRM_survey.
References in corpus (52)
- Sim-to-Real Transfer of Robotic Control with Dynamics Randomization
- DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation
- Vision-based Robotic Grasping From Object Localization, Object Pose Estimation to Grasp Estimation for Parallel Grippers: A Review
- More Than a Feeling: Learning to Grasp and Regrasp using Vision and Touch
- PointNetGPD: Detecting Grasp Configurations from Point Sets
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
- A Survey of Robot Manipulation in Contact
- TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors
- TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration
- Robot Cooking with Stir-fry: Bimanual Non-prehensile Manipulation of Semi-fluid Objects
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Deep Dynamics Models for Learning Dexterous Manipulation
- Assessing Transferability from Simulation to Reality for Reinforcement Learning
- Dex-NeRF: Using a Neural Radiance Field to Grasp Transparent Objects
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- How to select and use tools? : Active Perception of Target Objects Using Multimodal Deep Learning
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Towards Real-World Category-level Articulation Pose Estimation
- Exploiting Kinematic Redundancy for Robotic Grasping of Multiple Objects
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
- Grasp Multiple Objects with One Hand
- Tactile Tool Manipulation
- AllSight: A Low-Cost and High-Resolution Round Tactile Sensor with Zero-Shot Learning Capability
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- On the Role of the Action Space in Robot Manipulation Learning and Sim-to-Real Transfer
- Learning Arbitrary-Goal Fabric Folding with One Hour of Real Robot Experience
- Continuous Object State Recognition for Cooking Robots Using Pre-Trained Vision-Language Models and Black-box Optimization
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
- Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation Learning
- ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations
- Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play
- Deep Semantic Parsing of Freehand Sketches with Homogeneous Transformation, Soft-Weighted Loss, and Staged Learning
- SE(3)-Equivariant Relational Rearrangement with Neural Descriptor Fields
- RVT: Robotic View Transformer for 3D Object Manipulation
- Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated Objects
- Leveraging Language for Accelerated Learning of Tool Manipulation
- Co-GAIL: Learning Diverse Strategies for Human-Robot Collaboration
- Goal-Conditioned Imitation Learning using Score-based Diffusion Policies
- MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
- Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
- EDGI: Equivariant Diffusion for Planning with Embodied Agents
- SAM-RL: Sensing-Aware Model-Based Reinforcement Learning via Differentiable Physics-Based Simulation and Rendering
- Regularizing Model-Based Planning with Energy-Based Models
- DexDeform: Dexterous Deformable Object Manipulation with Human Demonstrations and Differentiable Physics
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation