Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
STAR: Rethinking MoE Routing as Structure-Aware Subspace Learning
Sumin Park, Noseong Park
Mixture-of-Experts (MoE) scales model capacity efficiently by selectively routing inputs to a specialized subset of experts. However, input-expert specialization, the core motivati…
cs.AI2026
Q-Delta: Beyond Key-Value Associative State Evolution
Sumin Park, Seojin Kim, Noseong Park
Linear attention reformulates sequence modeling as recurrent state evolution, enabling efficient linear-time inference. Under the key-value associative paradigm, existing approache…