2 papers
cs.AI2026
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Jinwei Kong, Runqi Meng, Fanyi Wang +4
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…
cs.CL2026
DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory
Wentao Qiu, Haotian Hu, Fanyi Wang +2
Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a fidelity--efficiency trade-…