Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Grouped Differential Attention
Junghwan Lim, Sungmin Lee, Dongseok Kim +7
The self-attention mechanism, while foundational to modern Transformer architectures, suffers from a critical inefficiency: it frequently allocates substantial attention to redunda…
cs.LG2025
Motif 2.6B Technical Report
Junghwan Lim, Sungmin Lee, Dongseok Kim +22
Recent advancements in Large Language Models (LLMs) have revolutionized artificial intelligence, yet developing an effective foundational LLM that balances high performance with co…
cs.LG2024
Effective Heterogeneous Federated Learning via Efficient Hypernetwork-based Weight Generation
Yujin Shin, Kichang Lee, Sungmin Lee +3
While federated learning leverages distributed client resources, it faces challenges due to heterogeneous client capabilities. This necessitates allocating models suited to clients…