2 papers
cs.LG2025
Fractional neural attention for efficient multiscale sequence processing
Cheng Kevin Qu, Andrew Ly, Pulin Gong
Attention mechanisms underpin the computational power of Transformer models, which have achieved remarkable success across diverse domains. Yet understanding and extending the prin…
cs.LG2025
Riddled basin geometry sets fundamental limits to predictability and reproducibility in deep learning
Andrew Ly, Pulin Gong
Fundamental limits to predictability are central to our understanding of many physical and computational systems. Here we show that, despite its remarkable capabilities, deep learn…