3 papers
cs.AI2026
Laguna M.1/XS.2 Technical Report
Julien Abadji, Marah Abdin, Connor Adams +93
We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has B total parameters (B activated per tok…
cs.LG2025
Understanding Differential Transformer Unchains Pretrained Self-Attentions
Chaerin Kong, Jiho Jang, Nojun Kwak
Differential Transformer has recently gained significant attention for its impressive empirical performance, often attributed to its ability to perform noise canceled attention. Ho…
cs.CV2024
Conservative Generator, Progressive Discriminator: Coordination of Adversaries in Few-shot Incremental Image Synthesis
Chaerin Kong, Nojun Kwak
The capacity to learn incrementally from an online stream of data is an envied trait of human learners, as deep neural networks typically suffer from catastrophic forgetting and st…