77 citations · 81 across the 3 of their papers we have counts for
3 papers
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
Matthias Seeger, Zeyu Zhang, Vihang Patil +2
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive…
Deep Explicit Duration Switching Models for Time Series
Abdul Fatir Ansari, Konstantinos Benidis, Richard Kurle +5
Many complex time series can be effectively subdivided into distinct regimes that exhibit persistent dynamics. Discovering the switching behavior and the statistical patterns in th…
GluonTS: Probabilistic Time Series Models in Python
Alexander Alexandrov, Konstantinos Benidis, Michael Bohlke-Schneider +10
We introduce Gluon Time Series (GluonTS, available at https://gluon-ts.mxnet.io), a library for deep-learning-based time series modeling. GluonTS simplifies the development of and…