Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
Muchen Li, Leonid Sigal, Renjie Liao
Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables…
cs.CL2025
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
Sadegh Mahdavi, Muchen Li, Kaiwen Liu +3
Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are cons…