3 papers
cs.AR2026
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
Zizhuo Fu, Yifan Zhou, Zhaoxin Lu +4
Multi-FPGA systems (MFS) are widely adopted for VLSI emulation and rapid prototyping. In an MFS, FPGAs connect only to a limited number of neighbors through bandwidth-constrained l…
cs.CR2026
Cachemir: Fully Homomorphic Encrypted Inference of Generative Large Language Model with KV Cache
Ye Yu, Yifan Zhou, Yi Chen +3
Generative large language models (LLMs) have revolutionized multiple domains. Modern LLMs predominantly rely on an autoregressive decoding strategy, which generates output tokens s…
cs.CR2025
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
Yifan Zhou, Tianshi Xu, Jue Hong +2
Private large language model (LLM) inference based on cryptographic primitives offers a promising path towards privacy-preserving deep learning. However, existing frameworks only s…