2 papers
cs.CL2025
Scale-invariant Attention
Ben Anson, Xi Wang, Laurence Aitchison
One persistent challenge in LLM research is the development of attention mechanisms that are able to generalise from training on shorter contexts to inference on longer contexts. W…
cs.LG2025
Revisiting Gradient Descent: A Dual-Weight Method for Improved Learning
Xi Wang
We introduce a novel framework for learning in neural networks by decomposing each neuron's weight vector into two distinct parts, and , thereby modeling contrastive inf…