Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection
Debarshi Kundu, Swaroop Ghosh, Vasant Honavar
Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive…
cs.LG2026
Projection-Free Transformers via Gaussian Kernel Attention
Debarshi Kundu, Archisman Ghosh, Swaroop Ghosh +1
Self-attention in Transformers is typically implemented as , where , , and are learned linear projections of the input…