From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention
Subham Singh, Ashutosh Mishra, Subha Raut
Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typicall…
cs.LG2026
Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps
Subham Singh, Ashutosh Mishra, Subha Raut
The paper analyzes when a cooldown phase in warmup‑stable‑decay learning‑rate schedules improves final loss, showing that the effect depends jointly on the structure of gradient no…