1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
How Can Mamba Learn In Context with Outliers and Generalize Provably?
Hongkang Li, Songtao Lu, Xiaodong Cui +2
The Mamba model has gained significant attention for its computational advantages over Transformer-based models, while achieving comparable performance across a wide range of langu…
Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
Xiaodong Cui, A F M Saif, Brian Kingsbury +1
Self-supervised pre-training using unlabeled data is widely used in automatic speech recognition. In this paper, we propose a new self-supervised pre-training approach to dealing w…
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
Hongkang Li, Songtao Lu, Pin-Yu Chen +2
Chain-of-Thought (CoT) is an efficient prompting method that enables the reasoning ability of large language models by augmenting the query using multiple examples with multiple in…
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
Hongkang Li, Meng Wang, Songtao Lu +2
Transformer-based large language models have displayed impressive in-context learning capabilities, where a pre-trained model can handle new tasks without fine-tuning by simply aug…