2 papers
cs.LG2026
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
Ruichen Xu, Kexin Chen
Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite growing empirical evidence that MLP lay…
cs.LG2025
Rethinking Benign Overfitting in Two-Layer Neural Networks
Ruichen Xu, Kexin Chen
Recent theoretical studies (Kou et al., 2023; Cao et al., 2022) have revealed a sharp phase transition from benign to harmful overfitting when the noise-to-feature ratio exceeds a…