Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
Chenwei Cui, Rockwell Jackson, Benjamin Joseph Herrera +2
Large language models have transformed many applications but remain expensive to train. Sparse Mixture of Experts (MoE) addresses this through conditional computation, with Expert…
cs.LG2024
An All-MLP Sequence Modeling Architecture That Excels at Copying
Chenwei Cui, Zehao Yan, Gedeon Muhawenayo +1
Recent work demonstrated Transformers' ability to efficiently copy strings of exponential sizes, distinguishing them from other architectures. We present the Causal Relation Networ…