activity
20242026
collaborators

5 papers

cs.LG2026

Unlocking Full Efficiency of Token Filtering in Large Language Model Training

Di Chai, Pengbo Li, Feiyuan Zhang +7

Token filtering has been proposed to enhance the utility of large language models (LLMs) by eliminating inconsequential tokens during training. While usingfewer tokens is expected…

cs.NI2025

MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training

Xudong Liao, Yijun Sun, Han Tian +13

Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dyn…

cs.NI2025

Swift: Rethinking RDMA Control Plane for Elastic Computing

Junxue Zhang, Han Tian, Xinyang Huang +5

Elastic computing enables dynamic scaling to meet workload demands, and Remote Direct Memory Access (RDMA) enhances this by providing high-throughput, low-latency network communica…

cs.AR2025

FLASH-FHE: A Heterogeneous Architecture for Fully Homomorphic Encryption Acceleration

Junxue Zhang, Xiaodian Cheng, Gang Cao +6

While many hardware accelerators have recently been proposed to address the inefficiency problem of fully homomorphic encryption (FHE) schemes, none of them is able to deliver opti…

cs.CR2024

PackVFL: Efficient HE Packing for Vertical Federated Learning

Liu Yang, Shuowei Cai, Di Chai +6

As an essential tool of secure distributed machine learning, vertical federated learning (VFL) based on homomorphic encryption (HE) suffers from severe efficiency problems due to d…