4 citations · 7 across the 14 of their papers we have counts for
14 papers
Dynamical Behaviors of the Gradient Flows for In-Context Learning
Songtao Lu, Yingdong Lu, Tomasz Nowicki
We derive the system of differential equations for the gradient flow characterizing the training process of linear in-context learning in full generality. Next, we explore the geom…
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
Shuchen Zhu, Boao Kong, Songtao Lu +2
This paper studies decentralized bilevel optimization, in which multiple agents collaborate to solve problems involving nested optimization structures with neighborhood communicati…
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
Xiaodong Cui, A F M Saif, Songtao Lu +4
In this paper, we propose a bilevel joint unsupervised and supervised training (BL-JUST) framework for automatic speech recognition. Compared to the conventional pre-training and f…
FADAS: Towards Federated Adaptive Asynchronous Optimization
Yujia Wang, Shiqiang Wang, Songtao Lu +1
Federated learning (FL) has emerged as a widely adopted training paradigm for privacy-preserving machine learning. While the SGD-based FL algorithms have demonstrated considerable…
Byzantine-Robust Decentralized Federated Learning
Minghong Fang, Zifan Zhang, Hairi +5
Federated learning (FL) enables multiple clients to collaboratively train machine learning models without revealing their private training data. In conventional FL, the system foll…
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
A F M Saif, Xiaodong Cui, Han Shen +3
In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term {bi-level joint…