Model-free Optical Processors using In Situ Reinforcement Learning with Proximal Policy Optimization
arXiv:2507.05583 · doi:10.1038/s41377-025-02148-7
Abstract
Optical computing holds promise for high-speed, energy-efficient information processing, with diffractive optical networks emerging as a flexible platform for implementing task-specific transformations. A challenge, however, is the effective optimization and alignment of the diffractive layers, which is hindered by the difficulty of accurately modeling physical systems with their inherent hardware imperfections, noise, and misalignments. While existing in situ optimization methods offer the advantage of direct training on the physical system without explicit system modeling, they are often limited by slow convergence and unstable performance due to inefficient use of limited measurement data. Here, we introduce a model-free reinforcement learning approach utilizing Proximal Policy Optimization (PPO) for the in situ training of diffractive optical processors. PPO efficiently reuses in situ measurement data and constrains policy updates to ensure more stable and faster convergence. We experimentally validated our method across a range of in situ learning tasks, including targeted energy focusing through a random diffuser, holographic image generation, aberration correction, and optical image classification, demonstrating in each task better convergence and performance. Our strategy operates directly on the physical system and naturally accounts for unknown real-world imperfections, eliminating the need for prior system knowledge or modeling. By enabling faster and more accurate training under realistic experimental constraints, this in situ reinforcement learning approach could offer a scalable framework for various optical and physical systems governed by complex, feedback-driven dynamics.
19 Pages, 7 Figures
References in corpus (22)
- Deep Learning with Coherent Nanophotonic Circuits
- All-Optical Machine Learning Using Diffractive Deep Neural Networks
- Deep physical neural networks enabled by a backpropagation algorithm for arbitrary physical systems
- Deep Image Prior
- All-optical Reservoir Computing
- Training of photonic neural networks through in situ backpropagation
- Reinforcement Learning in a large scale photonic Recurrent Neural Network
- Wave Physics as an Analog Recurrent Neural Network
- Experimentally realized in situ backpropagation for deep learning in nanophotonic neural networks
- Deep Learning with Coherent VCSEL Neural Networks
- Misalignment Resilient Diffractive Optical Networks
- The Forward-Forward Algorithm: Some Preliminary Investigations
- Trainable and Dynamic Computing: Error Backpropagation through Physical Media
- Optical Generative Models
- Physical deep learning based on optimal control of dynamical systems
- Wave-based extreme deep learning based on non-linear time-Floquet entanglement
- Roadmap on Neuromorphic Photonics
- Multiplexed gradient descent: Fast online training of modern datasets on hardware neural networks without backpropagation
- Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment
- High-performance real-world optical computing trained by in situ gradient-based model-free optimization
- Model-free front-to-end training of a large high performance laser neural network
- Physics-inspired Neuroacoustic Computing Based on Tunable Nonlinear Multiple-scattering