paper

Continual Policy Consolidation for Lifelong Robot Learning

arXiv:2601.22475

Abstract

Building a generalist robot policy requires continuously integrating new skills while preserving previously acquired behaviors. Directly optimizing a single policy over a growing task stream is difficult because robotic interaction is expensive, task distributions are heterogeneous, and sequential updates induce interference. To address these problems, we propose continual policy consolidation (CPC), a teacher--student framework that combines continual policy distillation with prioritized experience replay and expandable experts. This architecture separates skill acquisition from policy consolidation: teachers are trained independently through reinforcement learning, and their behaviors are continually distilled into a central generalist student. This decomposition retains the practical strength of reinforcement learning for task-specialized training while casting student-side consolidation as a supervised policy-learning problem. To balance stability and plasticity as the task stream grows, the student combines an expandable Transformer-based mixture-of-experts architecture with prioritized trajectory replay. Extensive experiments show that the student recovers a large proportion of teacher performance while achieving near-zero forgetting. These results demonstrate a scalable route for consolidating independently acquired robot skills into a continually growing generalist policy.

22 pages

Continual Policy Consolidation for Lifelong Robot Learning · wovepaper