4 papers
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
Hugo Hazard, Zafeirios Fountas, Martin A. Benfeghoul +3
Continual learning, one's ability to adapt to a sequence of tasks without forgetting previously acquired knowledge, remains a major challenge in machine learning and a key gap betw…
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
Frederico Wieser, Martin Benfeghoul, Haitham Bou Ammar +2
The rigid, uniform allocation of computation in standard Transformer (TF) architectures can limit their efficiency and scalability, particularly for large-scale models and long seq…
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
Martin Benfeghoul, Teresa Delgado, Adnan Oomerjee +3
Transformers' quadratic computational complexity limits their scalability despite remarkable performance. While linear attention reduces this to linear complexity, pre-training suc…
Human-inspired Episodic Memory for Infinite Context LLMs
Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee +4
Large language models (LLMs) have shown remarkable capabilities, but still struggle with processing extensive contexts, limiting their ability to maintain coherence and accuracy ov…