14 papers
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
Sajad Movahedi, Vera MilovanoviÄ, Shlomo Libo Feigin +5
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by loo…
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…
Low-Pass Flow Matching
Francesco M. Ruscio, T. Konstantin Rusch
Flow Matching typically relies on white noise sources, a choice often misaligned with the power spectra of natural data, which tend to decay with frequency. To address this, we int…
Neural Low-Discrepancy Sequences
Michael Etienne Van Huffel, Nathan Kirk, Makram Chahine +2
Low-discrepancy points are designed to efficiently fill the space in a uniform manner. This uniformity is highly advantageous in many problems in science and engineering, including…
The Curious Case of In-Training Compression of State Space Models
Makram Chahine, Philipp Nazari, Daniela Rus +1
State Space Models (SSMs), developed to tackle long sequence modeling tasks efficiently, offer both parallelizable training and fast inference. At their core are recurrent dynamica…
The Key to State Reduction in Linear Attention: A Rank-based Perspective
Philipp Nazari, T. Konstantin Rusch
Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate that the hidden state of trained lin…