2 papers
cs.LG2026
Mashup Learning: Faster Finetuning by Remixing Past Checkpoints
Sofia Maria Lo Cicero Vaina, Artem Chumachenko, Max Ryabinin
Finetuning on domain-specific data is a well-established method for enhancing LLM performance on downstream tasks. Training on each dataset produces a new set of model weights, res…
cs.LG2024
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
Yaroslav Aksenov, Nikita Balagansky, Sofia Maria Lo Cicero Vaina +3
Advancing the frontier of subquadratic architectures for Language Models (LMs) is crucial in the rapidly evolving field of natural language processing. Current innovations, includi…