3 papers
cs.CL2026
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
Difan Deng, Andreas Bentzen Winje, Lukas Fehring +1
The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios. In contrast, linear attention model families provide a promising d…
cs.CL2025
Neural Attention Search
Difan Deng, Marius Lindauer
We present Neural Attention Search (NAtS), a framework that automatically evaluates the importance of each token within a sequence and determines if the corresponding token can be…
cs.LG2025
Optimizing Time Series Forecasting Architectures: A Hierarchical Neural Architecture Search Approach
Difan Deng, Marius Lindauer
The rapid development of time series forecasting research has brought many deep learning-based modules in this field. However, despite the increasing amount of new forecasting arch…