2 papers
cs.AI2026
Laguna M.1/XS.2 Technical Report
Julien Abadji, Marah Abdin, Connor Adams +93
We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has B total parameters (B activated per tok…
cs.LG2025
Understanding Differential Transformer Unchains Pretrained Self-Attentions
Chaerin Kong, Jiho Jang, Nojun Kwak
Differential Transformer has recently gained significant attention for its impressive empirical performance, often attributed to its ability to perform noise canceled attention. Ho…