2 papers
cs.CL2026
Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
Matthias Seeger, Zeyu Zhang, Vihang Patil +2
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive…
stat.ML2024
Hyperparameter Optimization in Machine Learning
Luca Franceschi, Michele Donini, Valerio Perrone +5
Hyperparameters are configuration variables controlling the behavior of machine learning algorithms. They are ubiquitous in machine learning and artificial intelligence and the cho…