3 papers
cs.LG2026
Remember to Forget: Gated Adaptive Positional Encoding
Riccardo Ali, Alessio Borgi, Christopher Irwin +2
Rotary Positional Encoding (RoPE) is widely used in modern large language models. However, when sequences are extended beyond the range seen during training, rotary phases can ente…
cs.LG2026
Entropy-Lens: Uncovering Decision Strategies in LLMs
Riccardo Ali, Francesco Caso, Christopher Irwin +1
In large language models (LLMs), each block operates on the residual stream to map input token sequences to output token distributions. However, most of the interpretability litera…
cs.CL2025
HealthBranches: Synthesizing Clinically-Grounded Question Answering Datasets via Decision Pathways
Cristian Cosentino, Annamaria Defilippo, Marco Dossena +3
HealthBranches is a novel benchmark dataset for medical Question-Answering (Q&A), specifically designed to evaluate complex reasoning in Large Language Models (LLMs). This dataset…