4 papers
Delayed Attention Training Improves Length Generalization in Transformer--RNN Hybrids
Buu Phan, Reza Ebrahimi, Sanjay Haresh +1
We study length generalization in sequence models on a composite problem involving both state tracking and associative recall. Prior work finds that recurrent networks handle state…
Efficient Full-Stack Private Federated Deep Learning with Post-Quantum Security
Yiwei Zhang, Rouzbeh Behnia, Attila A. Yavuz +2
Federated learning (FL) enables collaborative model training while preserving user data privacy by keeping data local. Despite these advantages, FL remains vulnerable to privacy at…
Minimum Entropy Coupling with Bottleneck
M. Reza Ebrahimi, Jun Chen, Ashish Khisti
This paper investigates a novel lossy compression framework operating under logarithmic loss, designed to handle situations where the reconstruction distribution diverges from the…
Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers
MohammadReza Ebrahimi, Sunny Panchal, Roland Memisevic
Despite their recent successes, Transformer-based large language models show surprising failure modes. A well-known example of such failure modes is their inability to length-gener…