2 papers
cs.CR2026
FEnc: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment Encoding
Ran Ran, Zhaoting Gong, Nuo Xu +3
Fully Homomorphic Encryption (FHE) enables privacy-preserving machine learning but incurs extreme computational and memory overhead. These costs come not only from expensive low-le…
cs.LG2025
TASP: Topology-aware Sequence Parallelism
Yida Wang, Ke Hong, Xiuhong Li +4
Long-context large language models (LLMs) face constraints due to the quadratic complexity of the self-attention mechanism. The mainstream sequence parallelism (SP) method, Ring At…