1 paper
Vladimir Fedosov, Aleksandr Sazhin, Artemiy Grinenko +1
Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We combine Hierarchical Global Attention (HGA…