From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Token Geometry
Kathan Shah
The paper studies the distinct gradient geometry of token embedding and LM-head matrices in language models and introduces Ember, a lightweight optimizer that reduces memory usage…
cs.CV2026
Positional Embedding-Aware Activations
Kathan Shah, Chawin Sitawarin
We present a neural network architecture designed to naturally learn a positional embedding and overcome the spectral bias towards lower frequencies faced by conventional activatio…