2 papers
cs.CL2026
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
Narges Mokhtari, Farzan Haddadi, Ebrahim Rezaii
In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers…
cs.AI2026
Denoising Diffusion Generative Models Secretly Calculate Attentions
Farzan Haddadi, Leila Monfared, Ebrahim Rezaii +3
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer…