1k citations · 2.2k across the 32 of their papers we have counts for
Showing 2026 · cs.LGShow all
2 papers · 2 filters
cs.LG2026
Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Julien Siems, Riccardo Grazzi, Korbinian Pöppel +8
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that…
cs.LG2026
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
Niccolò Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan +4
We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates…