Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Remember to Forget: Gated Adaptive Positional Encoding
Riccardo Ali, Alessio Borgi, Christopher Irwin +2
Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during training, rotary phases can ente…
cs.LG2026
Entropy-Lens: Uncovering Decision Strategies in LLMs
Riccardo Ali, Francesco Caso, Christopher Irwin +1
In large language models (LLMs), each block operates on the residual stream to map input token sequences to output token distributions. However, most of the interpretability litera…