2 papers
cs.AR2026
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
Hanyuan Gao, Xiaoxuan Yang
Mixture-of-Experts (MoE) layers activate a subset of model weights, dubbed experts, to improve model performance. MoE is particularly promising for deployment on process-in-memory…
cs.LG2025
Norm-Q: Effective Compression Method for Hidden Markov Models in Neuro-Symbolic Applications
Hanyuan Gao, Xiaoxuan Yang
Hidden Markov models (HMM) are commonly used in generation tasks and have demonstrated strong capabilities in neuro-symbolic applications for the Markov property. These application…