4 papers
Decentralized SGD with Controlled Disagreement Finds Flatter Minima
Zesen Wang, Mikael Johansson
Decentralized training is often regarded as inferior to centralized training because the consensus errors between workers are thought to undermine convergence and generalization. T…
TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation
Zesen Wang, Lijuan Lan, Yonggang Li +1
Accurate carbon emission monitoring is critical for climate policy and emerging regulatory mechanisms such as the EU Carbon Border Adjustment Mechanism, yet city-level high-frequen…
YOTOnet: Zero-Shot Cross-Domain Fault Diagnosis via Domain-Conditioned Mixture of Experts
Zesen Wang, Zihao Wu, Yue Hu +2
Mechanical equipment forms the critical backbone of modern industrial production, yet domain shift severely limits the generalization of deep learning based fault diagnosis models…
From promise to practice: realizing high-performance decentralized training
Zesen Wang, Jiaojiao Zhang, Xuyang Wu +1
Decentralized training of deep neural networks has attracted significant attention for its theoretically superior scalability over synchronous data-parallel methods like All-Reduce…