Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Grouped Differential Attention
Junghwan Lim, Sungmin Lee, Dongseok Kim +7
The self-attention mechanism, while foundational to modern Transformer architectures, suffers from a critical inefficiency: it frequently allocates substantial attention to redunda…
cs.LG2025
Motif 2.6B Technical Report
Junghwan Lim, Sungmin Lee, Dongseok Kim +22
Recent advancements in Large Language Models (LLMs) have revolutionized artificial intelligence, yet developing an effective foundational LLM that balances high performance with co…