Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
Sumin Park, Noseong Park
Finding the optimal configuration of Sparse Mixture-ofExperts (SMoE) that maximizes semantic differentiation among experts is essential for exploiting the full potential of MoE arc…
cs.LG2025
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
Jeongwhan Choi, Seungjun Park, Sumin Park +2
Graph Neural Networks (GNNs) have emerged as powerful tools for learning on graph-structured data, but often struggle to balance local and global information. While graph Transform…