2 papers
cs.LG2026
Hasse Diagrams for Attention: A Partial Order Framework for Designing Transformer Masks
Chentao Li, Han Guo
During the training of large Transformer models, attention masks regulate the scope and direction of information flow across a sequence. Numerous mask variants exist, and operators…
cs.LG2025
Dual Refinement Cycle Learning: Unsupervised Text Classification of Mamba and Community Detection on Text Attributed Graph
Hong Wang, Yinglong Zhang, Hanhan Guo +2
Pretrained language models offer strong text understanding capabilities but remain difficult to deploy in real-world text-attributed networks due to their heavy dependence on label…