2 papers
cs.LG2026
Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging
Minsik Choi, Geewook Kim
Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered by gradient interference an…
cs.CL2026
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
Minsik Choi, Hyegang Son, Changhoon Kim +1
Transformer-based models have achieved remarkable performance in NLP tasks. However, their structural characteristics-multiple layers and attention heads-introduce efficiency chall…