1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2026
A Model with No Head and Many Thoughts
Nikita Koriagin, Yaroslav Aksenov, George Bredis +3
Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expr…
cs.LG2023★ 1 cited
Ahead-of-Time P-Tuning
Daniil Gavrilov, Nikita Balagansky
In this paper, we propose Ahead-of-Time (AoT) P-Tuning, a novel parameter-efficient fine-tuning method for pre-trained Language Models (LMs) that adds input-dependent bias before e…