4 papers
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
Frederico Wieser, Martin Benfeghoul, Haitham Bou Ammar +2
The rigid, uniform allocation of computation in standard Transformer (TF) architectures can limit their efficiency and scalability, particularly for large-scale models and long seq…
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
Martin Benfeghoul, Teresa Delgado, Adnan Oomerjee +3
Transformers' quadratic computational complexity limits their scalability despite remarkable performance. While linear attention reduces this to linear complexity, pre-training suc…
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
Xiaotong Ji, Shyam Sundhar Ramesh, Matthieu Zimmer +3
We introduce a novel inference-time alignment approach for LLMs that aims to generate safe responses almost surely, i.e., with probability approaching one. Our approach models the…
Efficient Reinforcement Learning with Large Language Model Priors
Xue Yan, Yan Song, Xidong Feng +4
In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases. However, they often require e…