Delocalized Photonic Deep Learning on the Internet's Edge
arXiv:2203.05466 · doi:10.1126/science.abq8271
Abstract
Advances in deep neural networks (DNNs) are transforming science and technology. However, the increasing computational demands of the most powerful DNNs limit deployment on low-power devices, such as smartphones and sensors -- and this trend is accelerated by the simultaneous move towards Internet-of-Things (IoT) devices. Numerous efforts are underway to lower power consumption, but a fundamental bottleneck remains due to energy consumption in matrix algebra, even for analog approaches including neuromorphic, analog memory and photonic meshes. Here we introduce and demonstrate a new approach that sharply reduces energy required for matrix algebra by doing away with weight memory access on edge devices, enabling orders of magnitude energy and latency reduction. At the core of our approach is a new concept that decentralizes the DNN for delocalized, optically accelerated matrix algebra on edge devices. Using a silicon photonic smart transceiver, we demonstrate experimentally that this scheme, termed Netcast, dramatically reduces energy consumption. We demonstrate operation in a photon-starved environment with 40 aJ/multiply of optical energy for 98.8% accurate image recognition and <1 photon/multiply using single photon detectors. Furthermore, we show realistic deployment of our system, classifying images with 3 THz of bandwidth over 86 km of deployed optical fiber in a Boston-area fiber network. Our approach enables computing on a new generation of edge devices with speeds comparable to modern digital electronics and power consumption that is orders of magnitude lower.
References in corpus (6)
- Array Programming with NumPy
- Attojoule Optoelectronics for Low-Energy Information Processing and Communications: a Tutorial Review
- Efficient, Compact and Low Loss Thermo-Optic Phase Shifter in Silicon
- An optical neural network using less than 1 photon per multiplication
- High-speed programmable photonic circuits in a cryogenically compatible, visible-NIR 200 mm CMOS architecture
- Post-Fabrication Trimming of Silicon Photonic Ring Resonators at Wafer-Scale
Cited by in corpus (22)
- The physics of optical computing
- Single chip photonic deep neural network with accelerated training
- Deep Learning with Coherent VCSEL Neural Networks
- Asymptotically Fault-Tolerant Programmable Photonics
- Photonic neural networks based on integrated silicon microresonators
- Training neural networks with end-to-end optical backpropagation
- Quantum-limited stochastic optical neural networks operating at a few quanta per activation
- Ultrafast single-channel machine vision based on neuro-inspired photonic computing
- Symmetric silicon microring resonator optical crossbar array for accelerated inference and training in deep learning
- All-optical nonlinear activation function based on stimulated Brillouin scattering
- An optoacoustic field-programmable perceptron for recurrent neural networks
- Waveguide-multiplexed photonic matrix-vector multiplication processor using multiport photodetectors
- Photonic crystal cavity IQ modulators in thin-film lithium niobate for coherent communications
- Synchronous micromechanically resonant programmable photonic circuits
- Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression
- High-performance real-world optical computing trained by in situ gradient-based model-free optimization
- Single-Shot Matrix-Matrix Multiplication Optical Tensor Processor for Deep Learning
- Architecture-Level Modeling of Photonic Deep Neural Network Accelerators
- Blending Optimal Control and Biologically Plausible Learning for Noise-Robust Physical Neural Networks
- A 262 TOPS Hyperdimensional Photonic AI Accelerator powered by a Si3N4 microcomb laser
- Scalable intensity-based photonic matrix-vector multiplication processor using single-wavelength time-division-multiplexed signals
- Quantum-secure multiparty deep learning