3 papers
cs.NI2026
Multi-stage Flow Scheduling for LLM Serving
Yijun Sun, Xudong Liao, Songrun Xie +5
Meeting stringent Time-To-First-Token (TTFT) requirements is crucial for LLM applications. To improve efficiency, modern LLM serving systems adopt disaggregated architectures with…
cs.NI2025
MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
Xudong Liao, Yijun Sun, Han Tian +13
Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dyn…
cs.AR2025
FLASH-FHE: A Heterogeneous Architecture for Fully Homomorphic Encryption Acceleration
Junxue Zhang, Xiaodian Cheng, Gang Cao +6
While many hardware accelerators have recently been proposed to address the inefficiency problem of fully homomorphic encryption (FHE) schemes, none of them is able to deliver opti…