3 papers
cs.LG2026
Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure
Xingjian Wang, Qingyu Han, Xiaodong Luo +1
Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate u…
cs.LG2026
Feature Augmentation of GNNs for ILPs: Local Uniqueness Suffices
Qingyu Han, Qian Li, Linxin Yang +3
Integer Linear Programs (ILPs) are central to real-world optimizations but notoriously difficult to solve. Learning to Optimize (L2O) has emerged as a promising paradigm, with Grap…
math.OC2025
SymILO: A Symmetry-Aware Learning Framework for Integer Linear Optimization
Qian Chen, Tianjian Zhang, Linxin Yang +5
Integer linear programs (ILPs) are commonly employed to model diverse practical problems such as scheduling and planning. Recently, machine learning techniques have been utilized t…