Learning to Synchronize in Minimum Time
arXiv:2608.10354
Abstract
The minimum-time feedback law for driving a population of coupled oscillators into synchrony is unknown. Here we settle it for identical Kuramoto oscillators under an instantaneous power constraint. Greedy control, maximizing at each instant, is exactly optimal at and suboptimal above, as dynamic programming confirms at . The obstruction is geometric: the greedy closed loop is a reparametrized gradient flow of , fixing its path independently of the power budget, and as a first-harmonic forcing it cannot leave a single Möbius orbit. A three-harmonic policy trained on the DP fields, a trajectory expert, and a smooth first-hitting-time objective beats greedy by -- at -- and matches DP to within where ground truth exists. Reading the policy rather than deploying it collapses it to a two-constant law, , which recovers -- of the network's advantage, with the same functional form holding from to . Machine learning discovered a closed-form law beyond the reach of direct analytical methods.