A Survey of Embodied Learning for Object-Centric Robotic Manipulation
arXiv:2408.11537 · doi:10.1007/s11633-025-1542-8
Abstract
Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike data-driven machine learning methods, embodied learning focuses on robot learning through physical interaction with the environment and perceptual feedback, making it especially suitable for robotic manipulation. In this paper, we provide a comprehensive survey of the latest advancements in this field and categorize the existing work into three main branches: 1) Embodied perceptual learning, which aims to predict object pose and affordance through various data representations; 2) Embodied policy learning, which focuses on generating optimal robotic decisions using methods such as reinforcement learning and imitation learning; 3) Embodied task-oriented learning, designed to optimize the robot's performance based on the characteristics of different tasks in object grasping and manipulation. In addition, we offer an overview and discussion of public datasets, evaluation metrics, representative applications, current challenges, and potential future research directions. A project associated with this survey has been established at https://github.com/RayYoh/OCRM_survey.
References in corpus (44)
- DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- A Survey of Robot Manipulation in Contact
- TACTO: A Fast, Flexible, and Open-source Simulator for High-Resolution Vision-based Tactile Sensors
- TransCG: A Large-Scale Real-World Dataset for Transparent Object Depth Completion and a Grasping Baseline
- Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration
- Robot Cooking with Stir-fry: Bimanual Non-prehensile Manipulation of Semi-fluid Objects
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Deep Dynamics Models for Learning Dexterous Manipulation
- Dex-NeRF: Using a Neural Radiance Field to Grasp Transparent Objects
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- How to select and use tools? : Active Perception of Target Objects Using Multimodal Deep Learning
- Eureka: Human-Level Reward Design via Coding Large Language Models
- Exploiting Kinematic Redundancy for Robotic Grasping of Multiple Objects
- Towards Real-World Category-level Articulation Pose Estimation
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
- Grasp Multiple Objects with One Hand
- Tactile Tool Manipulation
- AllSight: A Low-Cost and High-Resolution Round Tactile Sensor with Zero-Shot Learning Capability
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- On the Role of the Action Space in Robot Manipulation Learning and Sim-to-Real Transfer
- Learning Arbitrary-Goal Fabric Folding with One Hour of Real Robot Experience
- DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- Continuous Object State Recognition for Cooking Robots Using Pre-Trained Vision-Language Models and Black-box Optimization
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation Learning
- GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
- ObjectFolder: A Dataset of Objects with Implicit Visual, Auditory, and Tactile Representations
- Dexterity from Touch: Self-Supervised Pre-Training of Tactile Representations with Robotic Play
- SE(3)-Equivariant Relational Rearrangement with Neural Descriptor Fields
- RVT: Robotic View Transformer for 3D Object Manipulation
- Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated Objects
- Leveraging Language for Accelerated Learning of Tool Manipulation
- Co-GAIL: Learning Diverse Strategies for Human-Robot Collaboration
- MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
- Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
- Goal-Conditioned Imitation Learning using Score-based Diffusion Policies
- EDGI: Equivariant Diffusion for Planning with Embodied Agents
- Regularizing Model-Based Planning with Energy-Based Models
- SAM-RL: Sensing-Aware Model-Based Reinforcement Learning via Differentiable Physics-Based Simulation and Rendering
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation
- DexDeform: Dexterous Deformable Object Manipulation with Human Demonstrations and Differentiable Physics