collaborators

8 papers

hep-ph2026

Spontaneous Scoto-leptogenesis

Arghyajit Datta, Hyun Min Lee, Jun-Ho Song

We propose a low-scale spontaneous leptogenesis scenario within the dynamical minimal scotogenic model for accommodating neutrino masses and inert scalar dark matter simultaneously…

cs.AR2026

SpikON: A Dual-Parallel and Efficient Accelerator for Online Spiking Neural Networks Learning

Peilin Chen, Xiaoxuan Yang

Spiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient brain-inspired computing. However, existing online unsupervised SNN learning suffers from l…

cs.AR2026

Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching

Hanyuan Gao, Xiaoxuan Yang

Mixture-of-Experts (MoE) layers activate a subset of model weights, dubbed experts, to improve model performance. MoE is particularly promising for deployment on process-in-memory…

cs.AR2025

End-to-End Transformer Acceleration Through Processing-in-Memory Architectures

Xiaoxuan Yang, Peilin Chen, Tergel Molom-Ochir +1

Transformers have become central to natural language processing and large language models, but their deployment at scale faces three major challenges. First, the attention mechanis…

cs.LG2025

Norm-Q: Effective Compression Method for Hidden Markov Models in Neuro-Symbolic Applications

Hanyuan Gao, Xiaoxuan Yang

Hidden Markov models (HMM) are commonly used in generation tasks and have demonstrated strong capabilities in neuro-symbolic applications for the Markov property. These application…

cs.AR2025

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration

Peilin Chen, Xiaoxuan Yang

Large language models (LLMs) have gained great success in various domains. Existing systems cache Key and Value within the attention block to avoid redundant computations. However,…