Hybrid Reward-Driven Reinforcement Learning for Efficient Quantum Circuit Synthesis
arXiv:2507.16641 · doi:10.1007/s42484-026-00359-8
Abstract
A reinforcement learning (RL) framework is introduced for the efficient synthesis of quantum circuits that generate specified target quantum states from a fixed initial state, addressing a central challenge in both the Noisy Intermediate-Scale Quantum (NISQ) era and future fault-tolerant quantum computing. The approach utilizes tabular Q-learning, based on action sequences, within a discretized quantum state space, to effectively manage the exponential growth of the space dimension. The framework introduces a hybrid reward mechanism, combining a static, domain-informed reward that guides the agent toward the target state with customizable dynamic penalties that discourage inefficient circuit structures such as gate congestion and redundant state revisits. This is a circuit-aware reward, in contrast to the current trend of works on this topic, which are primarily fidelity-based. By leveraging sparse matrix representations and state-space discretization, the method enables practical navigation of high-dimensional environments while minimizing computational overhead. Benchmarking on graph-state preparation tasks for up to seven qubits, we demonstrate that the algorithm consistently discovers minimal-depth circuits with optimized gate counts. Moreover, extending the framework to a universal gate set still yields low depth circuits, highlighting the algorithm robustness and adaptability. The results confirm that this RL-driven approach, with our completely circuit-aware method, efficiently explores the complex quantum state space and synthesizes near-optimal quantum circuits, providing a resource-efficient foundation for quantum circuit optimization.
35 pages, 7 figures, color figures
References in corpus (28)
- Three qubits can be entangled in two inequivalent ways
- Measurement-based quantum computation with cluster states
- Multi-party entanglement in graph states
- Efficient quantum state tomography
- Quantum Computation as Geometry
- A meet-in-the-middle algorithm for fast synthesis of depth-optimal quantum circuits
- Information and Computation: Classical and Quantum Aspects
- A Survey of Multi-Objective Sequential Decision-Making
- Circuit complexity in quantum field theory
- Quantum error-correcting codes associated with graphs
- Reinforcement Learning in Different Phases of Quantum Control
- Active learning machine learns to create new quantum experiments
- Circuit complexity for free fermions
- Optimizing Quantum Error Correction Codes with Reinforcement Learning
- Time Evolution of Complexity: A Critique of Three Methods
- Models of quantum complexity growth
- A Study of Optimal 4-bit Reversible Toffoli Circuits and Their Synthesis
- Quantum complexity and topological phases of matter
- Quantum error correction for the toric code using deep reinforcement learning
- Global optimization of quantum dynamics with AlphaZero deep exploration
- Self-Correcting Quantum Many-Body Control using Reinforcement Learning with Tensor Networks
- Optimal preparation of graph states
- Integrability and complexity in quantum spin chains
- Unitary Synthesis of Clifford+T Circuits with Reinforcement Learning
- Reinforcement Learning Generation of 4-Qubits Entangled States
- A Substrate Scheduler for Compiling Arbitrary Fault-tolerant Graph States
- MLQM: Machine Learning Approach for Accelerating Optimal Qubit Mapping
- Towards Faster Reinforcement Learning of Quantum Circuit Optimization: Exponential Reward Functions