Real-time adaptive quantum error correction by model-free multi-agent learning
arXiv:2509.03974
The paper introduces a two-stage framework where multi-agent reinforcement learning discovers quantum error-correcting circuits offline, and a lightweight bandit-based retraining adapts these codes in real time to drifting noise, achieving significant reductions in logical infidelity.
Abstract
Quantum error correction (QEC) is essential for scalable quantum computing, yet existing approaches rely on static assumptions about noise that break down in realistic hardware, where error channels drift over time. We introduce a unified framework that separates QEC into two learning timescales: offline code discovery and online adaptation. Offline, Multi-Agent Reinforcement Learning (MARL) autonomously discovers complete QEC cycles as explicit quantum circuits, with separate agents responsible for encoding, syndrome extraction, and error recovery, and without prescribing a code family or circuit ansatz. Online, a lightweight adaptive layer, termed Bandit Retraining for Adaptive Variational Error Correction (BRAVE), continuously retunes a low-dimensional variational parameterization without retraining the full MARL stack. This yields a "discover once, adapt continuously" strategy that combines the flexibility of learned codes with real-time adaptation to non-stationary noise. At sufficiently high sampling rates relative to the noise drift, our method reduces logical infidelity by roughly 18-fold for qubit codes and 3-fold for qutrit codes compared to static error correction, while substantially extending robustness to noise fluctuations. These results establish a paradigm in which QEC is no longer static but is dynamically optimized for realistic quantum hardware.